Dualgauge Eval

Probes the joint functional correctness and security of LLM-generated code, while also evaluating an automated framework's ability to execute code in sandboxes and semantically judge test outcomes against human ground truth. Use when the user wants to benchmark on DualGauge-Bench, or asks about evaluating this task. Reports F1 Score.

qhjqhj00 cb9ec4e 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/dualgauge-eval commit cb9ec4e09a

Frequently asked questions

npx skillmds add qhjqhj00/dualgauge-eval