Loong Eval

Probes long-context multi-document question answering by requiring models to synthesize evidence from all provided documents (10K–250K+ tokens) across financial reports and academic papers. It tests information extraction, comparison, clustering, and chain-of-reasoning capabilities in heterogeneous, document-level agentic retrieval settings. Use when the user wants to benchmark on Loong, or asks about evaluating this task. Reports Avg Score.

qhjqhj00 67bdef2 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/loong-eval commit 67bdef297c

Frequently asked questions

npx skillmds add qhjqhj00/loong-eval