Mlissard Eval

This benchmark evaluates a model's ability to perform simple sequential reasoning tasks (e.g., counting, copying, list intersection) while extrapolating to longer input sequences. It specifically probes length generalization and the impact of multilingual in-context examples on reasoning robustness across different languages. Use when the user wants to benchmark on MLissard, or asks about evaluating this task. Reports accuracy.

qhjqhj00 4ae1ed8 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mlissard-eval commit 4ae1ed82fa

Frequently asked questions

npx skillmds add qhjqhj00/mlissard-eval