Sciqag 24d Eval

Evaluates open-ended, closed-book scientific question answering capabilities. It probes a model's ability to generate comprehensive, accurate, and reasonable answers to research-level science questions without external context or reference papers. Use when the user wants to benchmark on SciQAG-24D, SciQ, or asks about evaluating this task. Reports CAR.

qhjqhj00 519504a 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/sciqag-24d-eval commit 519504ab83

Frequently asked questions

npx skillmds add qhjqhj00/sciqag-24d-eval