Cweval Eval

Evaluates whether LLM-generated code is simultaneously functionally correct and secure against vulnerabilities. It measures the pass rate for functionality alone versus the joint pass rate for functionality and security, highlighting the gap where models produce working but vulnerable code. Use when the user wants to benchmark on CWEVAL-BENCH, or asks about evaluating this task. Reports func-sec@k.

qhjqhj00 321e27d 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/cweval-eval commit 321e27da23

Frequently asked questions

npx skillmds add qhjqhj00/cweval-eval