3 Secbench Large Scale Evaluation Suite Security

Evaluate and harden LLM-based autonomous agents against adversarial attacks using the α³-SecBench layered security framework. Assesses security (attack detection, CWE attribution), resilience (safe degradation), and trust (policy-compliant tool usage) across 7 autonomy layers. Use when: 'audit my LLM agent for security', 'add adversarial resilience to my autonomous system', 'evaluate agent trust and tool safety', 'harden my AI agent against prompt injection', 'security benchmark my LLM pipeline', 'test my agent for hallucinated tool calls'.

ndpvt-web 468ab79 15.6 KB Updated

File contents

ndpvt-web/arxiv-claude-skills/tree/main/skills/3-secbench-large-scale-evaluation-suite-security commit 468ab79dcc

Frequently asked questions

npx skillmds@latest add ndpvt-web/3-secbench-large-scale-evaluation-suite-security