Cyberseceval3 Human Eval

This evaluation probes the impact of LLM assistance on human cybersecurity practitioners' ability to execute novel cyberattack challenges. It measures objective performance metrics (phase completion rates and time) and subjective perception (sentiment/mental effort) across inexperienced and highly skilled cohorts, comparing LLM-assisted versus unassisted conditions. Use when the user wants to benchmark on CYBERSECEVAL 3 Challenge Set, or asks about evaluating this task. Reports phase completion time.

qhjqhj00 f0fbdc6 4.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/cyberseceval3-human-eval commit f0fbdc60d6

Frequently asked questions

npx skillmds add qhjqhj00/cyberseceval3-human-eval