Human Eval Functional Accuracy

Evaluates a code generation model's ability to produce correct, executable Python functions from docstrings and function signatures. It measures whether the generated code passes all provided unit tests for each programming problem. Use when the user wants to benchmark on HumanEval, or asks about evaluating this task. Reports functional accuracy.

qhjqhj00 dc6271b 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/human-eval-functional-accuracy commit dc6271b48f

Frequently asked questions

npx skillmds add qhjqhj00/human-eval-functional-accuracy