Codeflaws Repair Eval

This evaluation probes an LLM's ability to automatically detect and fix bugs in C programs by generating correct patches. It measures how effectively the model leverages test feedback, fault localization scores, and iterative reasoning to pass all provided test cases for each buggy submission. Use when the user wants to benchmark on Codeflaws, or asks about evaluating this task. Reports Repair Accuracy.

qhjqhj00 915a85d 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/codeflaws-repair-eval commit 915a85d289

Frequently asked questions

npx skillmds add qhjqhj00/codeflaws-repair-eval