Humaneval Mbpp Eval

Evaluates a model's ability to generate correct Python code for programming tasks and iteratively refine it using execution feedback or simulated human guidance. It measures both initial code generation quality and the effectiveness of a multi-turn debugging loop under strict runtime and edge-case constraints. Use when the user wants to benchmark on HumanEval, MBPP, HumanEval+, MBPP+, or asks about evaluating this task. Reports pass@1.

qhjqhj00 eec0b15 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/humaneval-mbpp-eval commit eec0b15535

Frequently asked questions

npx skillmds add qhjqhj00/humaneval-mbpp-eval