Recall Benchmark

Use when the user asks to measure the auditor's detection recall, run the recall benchmark, score a harness release against ground truth, bootstrap/admit benchmark targets, or asks "what fraction of real vulnerabilities do we find". Audits pinned pre-fix refs of repos with known-real findings (self-replay from the disposition ledger, CVE replay, seeded) on clean clones, scores detection with the finding_identity match ladder via match_benchmark.py, and emits benchmark.{json,md} + a metrics-ledger snapshot — recall overall/per-CWE/per-severity/per-language, held-out split, and run-to-run stability.

openshift 84e05d8 5 files · 48.3 KB Updated

File contents

openshift/traust/tree/main/harnessing/recall-benchmark commit 84e05d84e8

Frequently asked questions

npx skillmds@latest add openshift/recall-benchmark