Pencil Puzzle Bench Eval

Evaluates multi-step verifiable reasoning and agentic iteration on constraint-satisfaction puzzles. It probes a model's ability to plan, execute moves, check constraints step-by-step, and course-correct over long contexts. Use when the user wants to benchmark on Pencil Puzzle Bench, or asks about evaluating this task. Reports success rate.

qhjqhj00 afdd94d 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/pencil-puzzle-bench-eval commit afdd94ddee

Frequently asked questions

npx skillmds add qhjqhj00/pencil-puzzle-bench-eval