Slump Eval

Measures how much final code fidelity degrades when a system's design is progressively disclosed through multi-turn interaction rather than provided upfront. It evaluates semantic faithfulness to a committed design and structural integration of dependencies in long-horizon coding agents. Use when the user wants to benchmark on SLUMP benchmark, or asks about evaluating this task. Reports IF50.

qhjqhj00 d016dd3 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/slump-eval commit d016dd354a

Frequently asked questions

npx skillmds add qhjqhj00/slump-eval