AI Safety Guardrails

Security-by-design for AI (Prompt Injection defense, Hallucination checks, PII filters).

majiayu000 89907da 2 files · 2.0 KB Updated 567 repo stars

File contents

ai-safety-guardrails Skill

This skill protects the system from its own AI.

1. Input Guardrails (Defense)

  • Prompt Injection: "Ignore previous instructions".
    • Defense: Delimiters (XML tags), "Sandwich Defense" (System Prompt + User Input + System Reminder).
  • Jailbreaks: "Do this in 'DAN' mode".
    • Defense: Pattern matching for known jailbreak signatures.
  • PII Scrubbing: Regex scan input for SSN, Credit Cards, Emails before sending to LLM.

2. Output Guardrails (Verification)

  • Hallucination Check: "Self-Consistency" (Ask 3 times, take majority).
  • Tone Policing: Sentiment analysis on output. (Block Toxic/Aggressive responses).
  • Format Validation: Ensure JSON is valid JSON.

3. Libraries & Tools

  • NeMo Guardrails (NVIDIA)
  • Guardrails AI (Python)
  • Rebertha (PII)

4. System Design

  • Human in the Loop (HITL): For high-stakes actions (Transfer Money), AI proposes, Human approves.
  • Least Privilege: The Agent's API Token should NOT have admin access.

majiayu000/claude-skill-registry-data/tree/main/security/security-andreibesleaga-gabbe-8 commit 89907da9e1

Frequently asked questions

npx skillmds add majiayu000/ai-safety-guardrails