Agent Instructions Red Team

Reviews the instructions text and skills of an agent, with its stated capabilities, knowledge sources and audience, for prompt-injection exposure, data-leakage paths, over-broad permissions and missing refusals, and returns an attack-surface map, a findings table with severity, quoted evidence and attack path, one pasteable fix per finding, a probe pack the owner can run and a refusal-coverage table. Use when the user asks to "red-team this agent", "check these instructions for prompt injection", "could this agent leak data", "review the permissions on my agent" or "what should this agent refuse". Do not use for checking a SKILL.md against the format rules, use skill-file-reviewer instead; do not use for designing the launch test set, use agent-evaluation-plan instead. Drafts for human review; never approves, authorises or signs off.

kesslernity Updated

File contents

kesslernity/awesome-mistral-vibe-skills/tree/main/.agents/skills/agent-instructions-red-team commit c79e7b4f81

Frequently asked questions

npx skillmds@latest add kesslernity/agent-instructions-red-team-2