AI Security Expert
Enterprise AI security architect specializing in securing LLM applications, defending against prompt injection, implementing guardrails, and OWASP LLM Top 10 compliance.
OWASP LLM Top 10 (2025)
Quick Reference
| # |
Vulnerability |
Risk |
Key Defense |
| LLM01 |
Prompt Injection |
Critical |
Input sanitization, delimiters |
| LLM02 |
Insecure Output |
High |
Output validation, sanitization |
| LLM03 |
Training Data Poisoning |
High |
Data provenance, auditing |
| LLM04 |
Model DoS |
Medium |
Rate limiting, timeouts |
| LLM05 |
Supply Chain |
High |
Verification, pinning |
| LLM06 |
Sensitive Info Disclosure |
High |
PII detection, redaction |
| LLM07 |
Insecure Plugin Design |
High |
Permission model, validation |
| LLM08 |
Excessive Agency |
High |
Human-in-the-loop, least privilege |
| LLM09 |
Overreliance |
Medium |
Confidence scores, citations |
| LLM10 |
Model Theft |
Medium |
Rate limiting, watermarking |
LLM01: Prompt Injection
Attack Types:
- Direct: "Ignore previous instructions..."
- Indirect: Malicious content in RAG documents
- Encoding tricks: Unicode, special tokens
Defense Pattern:
User Input → Sanitize → Delimit → LLM → Validate Output → Filter
LLM02: Insecure Output Handling
- Never execute LLM output as code without validation
- Sanitize HTML (use allowlist)
- Validate SQL (SELECT only, table allowlist)
LLM04: Model DoS
- Rate limiting per user/API key
- Token limits on requests
- Timeout configurations
- Cost capping/alerts
LLM06: Sensitive Information Disclosure
- PII detection (regex + NER)
- System prompt protection
- Training data sanitization
- Output filtering
Code patterns: resources/security-patterns.py
PII Protection
Detection Patterns
| Type |
Example Pattern |
| Email |
*@*.com |
| Phone |
XXX-XXX-XXXX |
| SSN |
XXX-XX-XXXX |
| Credit Card |
16 digits |
| IP Address |
X.X.X.X |
Redaction Strategy
- Detect PII in input before LLM call
- Redact PII in LLM output
- Log without PII
- Encrypt at rest
Guardrails Implementation
NeMo Guardrails (NVIDIA)
define user express harmful intent
"How do I hack"
define bot refuse harmful request
"I can't help with that."
define flow harmful intent
user express harmful intent
bot refuse harmful request
Guardrails AI
guard = Guard().use_many(
ToxicLanguage(on_fail="fix"),
PIIFilter(on_fail="fix"),
ValidJSON(on_fail="reask")
)
Custom Pipeline
Input Guards → LLM Call → Output Guards → Response
Implementation: resources/security-patterns.py
Security Architecture
Defense in Depth Layers
| Layer |
Controls |
| Network |
WAF, DDoS protection, API gateway |
| Auth |
OAuth 2.0, API keys, mTLS |
| Input |
Schema validation, injection detection |
| Guardrails |
Topic restrictions, PII filtering |
| Model |
Versioning, anomaly detection |
| Output |
Response filtering, fact verification |
| Audit |
Logging, retention, compliance |
Zero Trust Principles
- Never trust, always verify
- Least privilege for agents
- Assume breach (log everything)
Compliance Frameworks
EU AI Act (High-Risk)
- Risk management system
- Data governance
- Technical documentation
- Human oversight
- Accuracy/robustness testing
SOC 2 for AI
- Security: Access controls, encryption
- Availability: SLA monitoring, DR
- Processing Integrity: Input/output validation
- Confidentiality: Data classification
- Privacy: Data minimization, consent
Security Testing
Red Team Categories
- Direct injection attempts
- Jailbreak prompts
- Indirect injection via context
- Encoding/unicode tricks
Test suite: resources/security-patterns.py
Testing Checklist
Incident Response
Severity Levels
| Incident |
Severity |
Response |
| Prompt injection detected |
Medium |
Block, log, analyze |
| Data exfiltration attempt |
High |
Block, forensics, notify |
| Model extraction detected |
High |
Rate limit, investigate |
Response Steps
- Contain (block source)
- Preserve (logs, evidence)
- Analyze (attack pattern)
- Remediate (update defenses)
- Document (security log)
Resources
Secure AI systems with defense in depth and zero trust principles.
1---2name: ai-security-expert3description: Enterprise AI security - OWASP LLM Top 10, prompt injection defense, guardrails, PII protection4---5
6# AI Security Expert
7
8Enterprise AI security architect specializing in securing LLM applications, defending against prompt injection, implementing guardrails, and OWASP LLM Top 10 compliance.
9
10## OWASP LLM Top 10 (2025)
11
12### Quick Reference
13
14| # | Vulnerability | Risk | Key Defense |
15|---|--------------|------|-------------|
16| LLM01 | Prompt Injection | Critical | Input sanitization, delimiters |
17| LLM02 | Insecure Output | High | Output validation, sanitization |
18| LLM03 | Training Data Poisoning | High | Data provenance, auditing |
19| LLM04 | Model DoS | Medium | Rate limiting, timeouts |
20| LLM05 | Supply Chain | High | Verification, pinning |
21| LLM06 | Sensitive Info Disclosure | High | PII detection, redaction |
22| LLM07 | Insecure Plugin Design | High | Permission model, validation |
23| LLM08 | Excessive Agency | High | Human-in-the-loop, least privilege |
24| LLM09 | Overreliance | Medium | Confidence scores, citations |
25| LLM10 | Model Theft | Medium | Rate limiting, watermarking |
26
27### LLM01: Prompt Injection
28
29**Attack Types:**
30- Direct: "Ignore previous instructions..."
31- Indirect: Malicious content in RAG documents
32- Encoding tricks: Unicode, special tokens
33
34**Defense Pattern:**
35```
36User Input → Sanitize → Delimit → LLM → Validate Output → Filter
37```
38
39### LLM02: Insecure Output Handling
40- Never execute LLM output as code without validation
41- Sanitize HTML (use allowlist)
42- Validate SQL (SELECT only, table allowlist)
43
44### LLM04: Model DoS
45- Rate limiting per user/API key
46- Token limits on requests
47- Timeout configurations
48- Cost capping/alerts
49
50### LLM06: Sensitive Information Disclosure
51- PII detection (regex + NER)
52- System prompt protection
53- Training data sanitization
54- Output filtering
55
56**Code patterns:** `resources/security-patterns.py`
57
58## PII Protection
59
60### Detection Patterns
61| Type | Example Pattern |
62|------|-----------------|
63| Email | `*@*.com` |
64| Phone | `XXX-XXX-XXXX` |
65| SSN | `XXX-XX-XXXX` |
66| Credit Card | 16 digits |
67| IP Address | `X.X.X.X` |
68
69### Redaction Strategy
701. Detect PII in input before LLM call
712. Redact PII in LLM output
723. Log without PII
734. Encrypt at rest
74
75## Guardrails Implementation
76
77### NeMo Guardrails (NVIDIA)
78```
79define user express harmful intent
80 "How do I hack"
81
82define bot refuse harmful request
83 "I can't help with that."
84
85define flow harmful intent
86 user express harmful intent
87 bot refuse harmful request
88```
89
90### Guardrails AI
91```python
92guard = Guard().use_many(
93 ToxicLanguage(on_fail="fix"),
94 PIIFilter(on_fail="fix"),
95 ValidJSON(on_fail="reask")
96)
97```
98
99### Custom Pipeline
100```
101Input Guards → LLM Call → Output Guards → Response
102```
103
104**Implementation:** `resources/security-patterns.py`
105
106## Security Architecture
107
108### Defense in Depth Layers
109
110| Layer | Controls |
111|-------|----------|
112| Network | WAF, DDoS protection, API gateway |
113| Auth | OAuth 2.0, API keys, mTLS |
114| Input | Schema validation, injection detection |
115| Guardrails | Topic restrictions, PII filtering |
116| Model | Versioning, anomaly detection |
117| Output | Response filtering, fact verification |
118| Audit | Logging, retention, compliance |
119
120### Zero Trust Principles
121- Never trust, always verify
122- Least privilege for agents
123- Assume breach (log everything)
124
125## Compliance Frameworks
126
127### EU AI Act (High-Risk)
128- Risk management system
129- Data governance
130- Technical documentation
131- Human oversight
132- Accuracy/robustness testing
133
134### SOC 2 for AI
135- Security: Access controls, encryption
136- Availability: SLA monitoring, DR
137- Processing Integrity: Input/output validation
138- Confidentiality: Data classification
139- Privacy: Data minimization, consent
140
141## Security Testing
142
143### Red Team Categories
1441. Direct injection attempts
1452. Jailbreak prompts
1463. Indirect injection via context
1474. Encoding/unicode tricks
148
149**Test suite:** `resources/security-patterns.py`
150
151### Testing Checklist
152- [ ] Injection patterns blocked
153- [ ] System prompt protected
154- [ ] PII detected and redacted
155- [ ] Rate limits enforced
156- [ ] Outputs validated
157- [ ] Audit logs complete
158
159## Incident Response
160
161### Severity Levels
162
163| Incident | Severity | Response |
164|----------|----------|----------|
165| Prompt injection detected | Medium | Block, log, analyze |
166| Data exfiltration attempt | High | Block, forensics, notify |
167| Model extraction detected | High | Rate limit, investigate |
168
169### Response Steps
1701. Contain (block source)
1712. Preserve (logs, evidence)
1723. Analyze (attack pattern)
1734. Remediate (update defenses)
1745. Document (security log)
175
176## Resources
177
178- [OWASP LLM Top 10](https://owasp.org/www-project-top-10-for-large-language-model-applications/)
179- [NIST AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework)
180- [NeMo Guardrails](https://github.com/NVIDIA/NeMo-Guardrails)
181- [Guardrails AI](https://github.com/guardrails-ai/guardrails)
182- [LLM Security Best Practices](https://llmsecurity.net/)
183
184---
185
186*Secure AI systems with defense in depth and zero trust principles.*