Ctibench Eval

Evaluates large language models on five cyber threat intelligence (CTI) tasks, including knowledge recall, vulnerability mapping, CVSS scoring, attack technique extraction, and threat attribution. It probes factual accuracy, logical reasoning, and contextual understanding within a domain-specific cybersecurity context. Use when the user wants to benchmark on CTIBench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 3c81068 5.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/ctibench-eval commit 3c810680c5

Frequently asked questions

npx skillmds add qhjqhj00/ctibench-eval