R1 Code Interpreter Eval

Probes LLMs' ability to autonomously generate and execute code for complex reasoning and planning tasks across logic, spatial, order, optimization, search, and math domains. It evaluates how well models can iteratively explore, optimize, and self-check solutions in multi-turn code execution environments. Use when the user wants to benchmark on SymBench, Big-Bench-Hard, Reasoning-Gym, or asks about evaluating this task. Reports exact match or constraint check.

qhjqhj00 d4d8430 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/r1-code-interpreter-eval commit d4d8430ea2

Frequently asked questions

npx skillmds add qhjqhj00/r1-code-interpreter-eval