Scitrek Eval

Evaluates long-context language models' ability to perform numerical aggregation, filtering, sorting, and logical operations across extended contexts (up to 1M tokens) using scientific article metadata and full-text articles. Use when the user wants to benchmark on SciTrek, or asks about evaluating this task. Reports exact match.

qhjqhj00 6ceeb9e 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/scitrek-eval commit 6ceeb9e6b6

Frequently asked questions

npx skillmds add qhjqhj00/scitrek-eval