Human Eval Cost Eval

This evaluation probes the accuracy-cost tradeoff of AI coding agents by measuring how often generated solutions pass test cases relative to the actual inference cost required. It highlights whether complex agent architectures provide genuine performance gains over simple retry baselines when compute expenses are accounted for. Use when the user wants to benchmark on HumanEval, or asks about evaluating this task. Reports accuracy.

qhjqhj00 357e3b9 2.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/human-eval-cost-eval commit 357e3b9c07

Frequently asked questions

npx skillmds add qhjqhj00/human-eval-cost-eval