Quality Eval

This benchmark evaluates a model's ability to comprehend and reason over long documents (2k–8k tokens) to answer multiple-choice questions. It specifically probes whether models can integrate global context rather than relying on local keyword matching or summaries, with a subset (HARD) filtering for questions that require full reading rather than skimming. Use when the user wants to benchmark on QuALITY, or asks about evaluating this task. Reports accuracy.

qhjqhj00 8478b5f 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/quality-eval commit 8478b5fcd5

Frequently asked questions

npx skillmds add qhjqhj00/quality-eval