Scholarqa Bench Eval

Evaluates LLMs' ability to synthesize scientific literature by answering open-ended, multi-domain questions using retrieved papers. It probes long-form generation, factual correctness, citation accuracy, and content quality/organization across single- and multi-paper retrieval setups. Use when the user wants to benchmark on ScholarQABench, or asks about evaluating this task. Reports Citation F1.

qhjqhj00 2d6cad0 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/scholarqa-bench-eval commit 2d6cad0f27

Frequently asked questions

npx skillmds add qhjqhj00/scholarqa-bench-eval