Scrolls Eval

Evaluates long-text understanding capabilities across summarization, question answering, and natural language inference tasks. It probes whether models can effectively process and extract information from documents exceeding standard context windows (up to 16K tokens) using chunked encoding and cross-chunk fusion. Use when the user wants to benchmark on SCROLLS, or asks about evaluating this task. Reports Avg SCROLLS score.

qhjqhj00 8e2d508 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/scrolls-eval commit 8e2d508c43

Frequently asked questions

npx skillmds add qhjqhj00/scrolls-eval