Scalar Eval

Evaluates long-context academic reasoning by testing whether LLMs can correctly identify masked citations within scientific papers. It probes the model's ability to understand semantic context, attributional claims, and descriptive references across varying context lengths and difficulty levels. Use when the user wants to benchmark on SCALAR, or asks about evaluating this task. Reports accuracy.

qhjqhj00 231fe14 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/scalar-eval commit 231fe14568

Frequently asked questions

npx skillmds add qhjqhj00/scalar-eval