Adversarial Agent Review

Gate an agent on a red-team suite that tries to make it misbehave — jailbreaks, injected instructions, scope escalations, harmful requests — each with the safe behavior it must show, scored objectively before it ships. Use this when an agent is exposed to adversarial input or acts with real consequences, when you need a repeatable security regression test rather than ad-hoc probing, or when "it seems safe" isn't evidence. Distinct from quality review and completeness checks: this asks whether the agent can be *made to fail*, not whether its output is good.

sharp-skills Updated

File contents

sharp-skills/skills/tree/main/skills/adversarial-agent-review commit 7a57877a96

Frequently asked questions

npx skillmds@latest add sharp-skills/adversarial-agent-review