Testing Agents For Indirect Prompt Injection

Test whether an AI agent obeys instructions hidden in the content it ingests, rather than only the user's. Enumerate every channel through which untrusted content reaches the model context (retrieved docs, fetched pages, uploaded files, emails, tool outputs, filenames, images and PDFs, other agents), plant channel-appropriate payloads, and measure whether they change the agent's actions. Use when reviewing any agent or LLM app that reads external content and can act. Covers channel enumeration, overt and covert payloads, canary observables, and impact via the trifecta.

UnboundCompute Updated

File contents

UnboundCompute/security-agent-skills/tree/main/skills/testing-agents-for-indirect-prompt-injection commit 112ea3e461

Frequently asked questions

npx skillmds@latest add unboundcompute/testing-agents-for-indirect-prompt-injection