Mus Eval

Evaluates large language models' ability to perform multi-step, commonsense-rich reasoning over long natural language narratives. It probes whether models can follow complex, implicit logical chains (e.g., murder motives, object spatial reasoning, team skill matching) without relying on simple keyword heuristics or rule-based shortcuts. Use when the user wants to benchmark on MuSR, or asks about evaluating this task. Reports accuracy.

qhjqhj00 334ddb9 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mus-eval commit 334ddb9253

Frequently asked questions

npx skillmds add qhjqhj00/mus-eval