Secure Eval

Evaluates large language models' capabilities in cybersecurity, specifically focusing on Industrial Control Systems (ICS). It probes knowledge extraction, vulnerability understanding, out-of-distribution reasoning, and risk evaluation using real-world threat intelligence sources. Use when the user wants to benchmark on MAET, CWET, KCV, VOOD, RERT, CPST, or asks about evaluating this task. Reports accuracy.

qhjqhj00 060438b 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/secure-eval commit 060438bee1

Frequently asked questions

npx skillmds add qhjqhj00/secure-eval