Lingoly Eval

This benchmark evaluates large language models' ability to perform multi-step linguistic reasoning and deductive puzzle solving in low-resource and extinct languages. It probes out-of-domain grammatical inference and instruction-following under conditions of minimal pre-training exposure, requiring models to extract and apply novel rules from provided context rather than relying on memorized knowledge. Use when the user wants to benchmark on LINGOLY, or asks about evaluating this task. Reports Exact Match.

qhjqhj00 af9f954 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/lingoly-eval commit af9f95469f

Frequently asked questions

npx skillmds add qhjqhj00/lingoly-eval