Humaneval Eval

Evaluates a model's ability to generate correct, executable Python code from natural language function descriptions and signatures. It measures functional correctness by checking if generated code passes hidden unit tests. Use when the user wants to benchmark on HumanEval, or asks about evaluating this task. Reports pass@1.

qhjqhj00 434f6fc 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/humaneval-eval commit 434f6fc06b

Frequently asked questions

npx skillmds add qhjqhj00/humaneval-eval