Opensec Eval

Probes incident response agent calibration under adversarial prompt injection. It measures how well models distinguish true threats from false positives, resist injection attacks, and execute containment actions without indiscriminately exhausting the available action space. Use when the user wants to benchmark on OpenSec Standard-Tier Episodes, or asks about evaluating this task. Reports Containment rate.

qhjqhj00 4c5f036 4.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/opensec-eval commit 4c5f036d7c

Frequently asked questions

npx skillmds add qhjqhj00/opensec-eval