Prompt Injection Denylist

Detect and investigate fraud rings that use Warp as a free LLM backend by injecting third-party system prompts (e.g., OpenClaw, Claude Code) into agent mode queries. Use when you see unusual agent mode cost spikes, high volumes of identical-looking first messages, or tool-call loops referencing tools that don't exist in Warp (e.g., Bash, Read, exec). This skill covers how to identify prompt injection patterns, backtest proposed denylist strings for false positives, and monitor cost impact.

warpdotdev bebe44e 10.3 KB Updated

File contents

warpdotdev/fraud-detection-agent-oss/tree/main/.agents/skills/prompt-injection-denylist commit bebe44ec49

Frequently asked questions

npx skillmds@latest add warpdotdev/prompt-injection-denylist