Secbench Eval

Evaluates large language models' cybersecurity knowledge retention and logical reasoning capabilities across multiple subdomains, languages, and difficulty levels using multiple-choice and short-answer questions. Use when the user wants to benchmark on SecBench, or asks about evaluating this task. Reports correctness percentage.

qhjqhj00 d34f0b2 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/secbench-eval commit d34f0b24e9

Frequently asked questions

npx skillmds add qhjqhj00/secbench-eval