Longbench Eval

Evaluates large language models' ability to understand and process long contexts across bilingual (English and Chinese) multitask scenarios, including single/multi-document QA, summarization, few-shot learning, code completion, and synthetic tasks. Use when the user wants to benchmark on LongBench, or asks about evaluating this task. Reports F1.

qhjqhj00 cef1499 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/longbench-eval commit cef14999c9

Frequently asked questions

npx skillmds add qhjqhj00/longbench-eval