Mmlu Sr Eval

This benchmark stress-tests the reasoning capability of large language models by replacing key terms in multiple-choice questions and answers with arbitrary dummy words and their definitions. It probes whether models rely on genuine conceptual understanding or merely on lexical memorization of pre-trained vocabulary. Performance is measured across three substitution variants to isolate the impact of context modification on reasoning robustness. Use when the user wants to benchmark on MMLU-SR, or asks about evaluating this task. Reports accuracy.

qhjqhj00 40a1897 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mmlu-sr-eval commit 40a1897a1d

Frequently asked questions

npx skillmds add qhjqhj00/mmlu-sr-eval