MLOps & ML Security - Complete Reference (Jan 2026)
Production ML lifecycle with modern security practices.
This skill covers:
- Production: Data ingestion, deployment, drift detection, monitoring, incident response
- Security: Prompt injection, jailbreak defense, RAG security, output filtering
- Governance: Privacy protection, supply chain security, safety evaluation
- Data ingestion (dlt): Load data from APIs, databases to warehouses
- Model deployment: Batch jobs, real-time APIs, hybrid systems, event-driven automation
- Operations: Real-time monitoring, drift detection, automated retraining, incident response
Modern Best Practices (Jan 2026):
- Version everything that can change: model artifacts, data snapshots, feature definitions, prompts/configs, and agent graphs; require reproducibility, rollbacks, and audit logs (NIST SSDF: https://csrc.nist.gov/pubs/sp/800/218/final).
- Gate changes with evals (offline + online) and safe rollout (shadow/canary/blue-green); treat regressions in quality, safety, latency, and cost as release blockers.
- Align controls and documentation to risk posture (EU AI Act: https://eur-lex.europa.eu/eli/reg/2024/1689/oj; NIST AI RMF + GenAI profile: https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf, https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf).
- Operationalize security: threat model the full system (data, model, prompts, tools, RAG), harden the supply chain (SBOM/signing), and ship incident playbooks for both reliability and safety events.
It is execution-focused:
- Data ingestion patterns (REST APIs, database replication, incremental loading)
- Deployment patterns (batch, online, hybrid, streaming, event-driven)
- Automated monitoring with real-time drift detection
- Automated retraining pipelines (monitor → detect → trigger → validate → deploy)
- Incident handling with validated rollback and postmortems
- Links to copy-paste templates in
assets/
Quick Reference
| Task |
Tool/Framework |
Command |
When to Use |
| Data Ingestion |
dlt (data load tool) |
dlt pipeline run, dlt init |
Loading from APIs, databases to warehouses |
| Batch Deployment |
Airflow, Dagster, Prefect |
airflow dags trigger, dagster job launch |
Scheduled predictions on large datasets |
| API Deployment |
FastAPI, Flask, TorchServe |
uvicorn app:app, torchserve --start |
Real-time inference (<500ms latency) |
| LLM Serving |
vLLM, TGI, BentoML |
vllm serve model, bentoml serve |
High-throughput LLM inference |
| Model Registry |
MLflow, W&B, ZenML |
mlflow.register_model(), zenml model register |
Versioning and promoting models |
| Drift Detection |
Statistical tests + monitors |
PSI/KS, embedding drift, prediction drift |
Detect data/process changes and trigger review |
| Monitoring |
Prometheus, Grafana |
prometheus.yml, Grafana dashboards |
Metrics, alerts, SLO tracking |
| AgentOps |
AgentOps, Langfuse, LangSmith |
agentops.init(), trace visualization |
AI agent observability, session replay |
| Incident Response |
Runbooks, PagerDuty |
Documented playbooks, alert routing |
Handling failures and degradation |
Use This Skill When
Use this skill when the user asks for deployment, operations, monitoring, incident handling, or governance for ML/LLM/agent systems, e.g.:
- "How do I deploy this model to prod?"
- "Design a batch + online scoring architecture."
- "Add monitoring and drift detection to our model."
- "Write an incident runbook for this ML service."
- "Package this LLM/RAG pipeline as an API."
- "Plan our retraining and promotion workflow."
- "Load data from Stripe API to Snowflake."
- "Set up incremental database replication with dlt."
- "Build an ELT pipeline for warehouse loading."
If the user is asking only about EDA, modelling, or theory, prefer:
ai-ml-data-science (EDA, features, modelling, SQL transformation with SQLMesh)
ai-llm (prompting, fine-tuning, eval)
ai-rag (retrieval pipeline design)
ai-llm-inference (compression, spec decode, serving internals)
If the user is asking about SQL transformation (after data is loaded), prefer:
ai-ml-data-science (SQLMesh templates for staging, intermediate, marts layers)
Decision Tree: Choosing Deployment Strategy
User needs to deploy: [ML System]
├─ Data Ingestion?
│ ├─ From REST APIs? → dlt REST API templates
│ ├─ From databases? → dlt database sources (PostgreSQL, MySQL, MongoDB)
│ └─ Incremental loading? → dlt incremental patterns (timestamp, ID-based)
│
├─ Model Serving?
│ ├─ Latency <500ms? → FastAPI real-time API
│ ├─ Batch predictions? → Airflow/Dagster batch pipeline
│ └─ Mix of both? → Hybrid (batch features + online scoring)
│
├─ Monitoring & Ops?
│ ├─ Drift detection? → Evidently + automated retraining triggers
│ ├─ Performance tracking? → Prometheus + Grafana dashboards
│ └─ Incident response? → Runbooks + PagerDuty alerts
│
└─ LLM/RAG Production?
├─ Cost optimization? → Caching, prompt templates, token budgets
└─ Safety? → See ai-mlops skill
Core Concepts (Vendor-Agnostic)
- Lifecycle loop: train → validate → deploy → monitor → respond → retrain/retire.
- Risk controls: access control, data minimization, logging, and change management (NIST AI RMF: https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf).
- Observability planes: system metrics (latency/errors), data metrics (freshness/drift), quality metrics (model performance).
- Incident readiness: detection, containment, rollback, and root-cause analysis.
Do / Avoid
Do
- Do gate deployments with repeatable checks: evaluation pass, load test, security review, rollback plan.
- Do version everything: code, data, features, model artifact, prompt templates, configuration.
- Do define SLOs and budgets (latency/cost/error rate) before optimizing.
Avoid
- Avoid manual “clickops” deployments without audit trail.
- Avoid silent upgrades; require eval + canary for model/prompt changes.
- Avoid drift dashboards without actions; every alert needs an owner and runbook.
Core Patterns Overview
This skill provides production-ready patterns and guides organized into comprehensive references:
Data & Infrastructure Patterns
Pattern 0: Data Contracts, Ingestion & Lineage
→ See Data Ingestion Patterns
- Data contracts with SLAs and versioning
- Ingestion modes (CDC, batch, streaming)
- Lineage tracking and schema evolution
- Replay and backfill procedures
Pattern 1: Choose Deployment Mode
→ See Deployment Patterns
- Decision table (batch, online, hybrid, streaming)
- When to use each mode
- Deployment mode selection checklist
Pattern 2: Standard Deployment Lifecycle
→ See Deployment Lifecycle
- Pre-deploy, deploy, observe, operate, evolve phases
- Environment promotion (dev → staging → prod)
- Gradual rollout strategies (canary, blue-green)
Pattern 3: Packaging & Model Registry
→ See Model Registry Patterns
- Model registry structure and metadata
- Packaging strategies (Docker, ONNX, MLflow)
- Promotion flows (experimental → production)
- Versioning and governance
Serving Patterns
Pattern 4: Batch Scoring Pipeline
→ See Deployment Patterns
- Orchestration with Airflow/Dagster
- Idempotent scoring jobs
- Validation and backfill procedures
Pattern 5: Real-Time API Scoring
→ See API Design Patterns
- Service design (HTTP/JSON, gRPC)
- Input/output schemas
- Rate limiting, timeouts, circuit breakers
Pattern 6: Hybrid & Feature Store Integration
→ See Feature Store Patterns
- Batch vs online features
- Feature store architecture
- Training-serving consistency
- Point-in-time correctness
Operations Patterns
Pattern 7: Monitoring & Alerting
→ See Monitoring Best Practices
- Data, performance, and technical metrics
- SLO definition and tracking
- Dashboard design and alerting strategies
Pattern 8: Drift Detection & Automated Retraining
→ See Drift Detection Guide
- Automated retraining triggers
- Event-driven retraining pipelines
Pattern 9: Incidents & Runbooks
→ See Incident Response Playbooks
- Common failure modes
- Detection, diagnosis, resolution
- Post-mortem procedures
Pattern 10: LLM / RAG in Production
→ See LLM & RAG Production Patterns
- Prompt and configuration management
- Safety and compliance (PII, jailbreaks)
- Cost optimization (token budgets, caching)
- Monitoring and fallbacks
Pattern 11: Cross-Region, Residency & Rollback
→ See Multi-Region Patterns
- Multi-region deployment architectures
- Data residency and tenant isolation
- Disaster recovery and failover
- Regional rollback procedures
Pattern 12: Online Evaluation & Feedback Loops
→ See Online Evaluation Patterns
- Feedback signal collection (implicit, explicit)
- Shadow and canary deployments
- A/B testing with statistical significance
- Human-in-the-loop labeling
- Automated retraining cadence
Pattern 13: AgentOps (AI Agent Operations)
→ See AgentOps Patterns
- Session tracing and replay for AI agents
- Cost and latency tracking across agent runs
- Multi-agent visualization and debugging
- Tool invocation monitoring
- Integration with CrewAI, LangGraph, OpenAI Agents SDK
Pattern 14: Edge MLOps & TinyML
→ See Edge MLOps Patterns
- Device-aware CI/CD pipelines
- OTA model updates with rollback
- Federated learning operations
- Edge drift detection
- Intermittent connectivity handling
Resources (Detailed Guides)
For comprehensive operational guides, see:
Core Infrastructure:
- Data Ingestion Patterns - Data contracts, CDC, batch/streaming ingestion, lineage, schema evolution
- Deployment Lifecycle - Pre-deploy validation, environment promotion, gradual rollout, rollback
- Model Registry Patterns - Versioning, packaging, promotion workflows, governance
- Feature Store Patterns - Batch/online features, hybrid architectures, consistency, latency optimization
Serving & APIs:
- Deployment Patterns - Batch, online, hybrid, streaming deployment strategies and architectures
- API Design Patterns - ML/LLM/RAG API patterns, input/output schemas, reliability patterns, versioning
Operations & Reliability:
- Monitoring Best Practices - Metrics collection, alerting strategies, SLO definition, dashboard design
- Drift Detection Guide - Statistical tests, automated detection, retraining triggers, recovery strategies
- Incident Response Playbooks - Runbooks for common failure modes, diagnostics, resolution steps
Security & Governance:
- Threat Models - Trust boundaries, attack surface, control mapping
- Prompt Injection Mitigation - Input hardening, tool/RAG containment, least privilege
- Jailbreak Defense - Robust refusal behavior, safe completion patterns
- RAG Security - Retrieval poisoning, context injection, sensitive data leakage
- Output Filtering - Layered filters (PII/toxicity/policy), block/rewrite strategies
- Privacy Protection - PII handling, data minimization, retention, consent
- Supply Chain Security - SBOM, dependency pinning, artifact signing
- Safety Evaluation - Red teaming, eval sets, incident readiness
Advanced Patterns:
- LLM & RAG Production Patterns - Prompt management, safety, cost optimization, caching, monitoring
- Multi-Region Patterns - Multi-region deployment, data residency, disaster recovery, rollback
- Online Evaluation Patterns - A/B testing, shadow deployments, feedback loops, automated retraining
- AgentOps Patterns - AI agent observability, session replay, cost tracking, multi-agent debugging
- Edge MLOps Patterns - TinyML, federated learning, OTA updates, device-aware CI/CD
Templates
Use these as copy-paste starting points for production artifacts:
Data Ingestion (dlt)
For loading data into warehouses and pipelines:
- dlt basic pipeline setup - Install, configure, run basic extraction and loading
- dlt REST API sources - Extract from REST APIs with pagination, authentication, rate limiting
- dlt database sources - Replicate from PostgreSQL, MySQL, MongoDB, SQL Server
- dlt incremental loading - Timestamp-based, ID-based, merge/upsert patterns, lookback windows
- dlt warehouse loading - Load to Snowflake, BigQuery, Redshift, Postgres, DuckDB
Use dlt when:
- Loading data from APIs (Stripe, HubSpot, Shopify, custom APIs)
- Replicating databases to warehouses
- Building ELT pipelines with incremental loading
- Managing data ingestion with Python
For SQL transformation (after ingestion), use:
→ ai-ml-data-science skill (SQLMesh templates for staging/intermediate/marts layers)
Deployment & Packaging
- Deployment & MLOps template - Complete MLOps lifecycle, model registry, promotion workflows
- Deployment readiness checklist - Go/No-Go gate, monitoring, and rollback plan
- API service template - Real-time REST/gRPC API with FastAPI, input validation, rate limiting
- Batch scoring pipeline template - Orchestrated batch inference with Airflow/Dagster, validation, backfill
Monitoring & Operations
- Monitoring & alerting template - Data/performance/technical metrics, dashboards, SLO definition
- Drift detection & retraining template - Automated drift detection, retraining triggers, promotion pipelines
- Incident runbook template - Failure mode playbooks, diagnosis steps, resolution procedures
Navigation
Resources
- references/drift-detection-guide.md
- references/model-registry-patterns.md
- references/online-evaluation-patterns.md
- references/monitoring-best-practices.md
- references/llm-rag-production-patterns.md
- references/api-design-patterns.md
- references/incident-response-playbooks.md
- references/deployment-patterns.md
- references/data-ingestion-patterns.md
- references/deployment-lifecycle.md
- references/feature-store-patterns.md
- references/multi-region-patterns.md
- references/agentops-patterns.md
- references/edge-mlops-patterns.md
Templates
Data
- data/sources.json - Curated external references
External Resources
See data/sources.json for curated references on:
- Serving frameworks (FastAPI, Flask, gRPC, TorchServe, KServe, Ray Serve)
- Orchestration (Airflow, Dagster, Prefect)
- Model registries and MLOps (MLflow, W&B, Vertex AI, Sagemaker)
- Monitoring and observability (Prometheus, Grafana, OpenTelemetry, Evidently)
- Feature stores (Feast, Tecton, Vertex, Databricks)
- Streaming & messaging (Kafka, Pulsar, Kinesis)
- LLMOps & RAG infra (vector DBs, LLM gateways, safety tools)
Data Lake & Lakehouse
For comprehensive data lake/lakehouse patterns (beyond dlt ingestion), see data-lake-platform:
- Table formats: Apache Iceberg, Delta Lake, Apache Hudi
- Query engines: ClickHouse, DuckDB, Apache Doris, StarRocks
- Alternative ingestion: Airbyte (GUI-based connectors)
- Transformation: dbt (alternative to SQLMesh)
- Streaming: Apache Kafka patterns
- Orchestration: Dagster, Airflow
This skill focuses on ML-specific deployment, monitoring, and security. Use data-lake-platform for general-purpose data infrastructure.
Recency Protocol (Tooling Recommendations)
When users ask recommendation questions about MLOps tooling, verify recency before answering.
Trigger Conditions
- "What's the best MLOps platform for [use case]?"
- "What should I use for [deployment/monitoring/drift detection]?"
- "What's the latest in MLOps?"
- "Current best practices for [model registry/feature store/observability]?"
- "Is [MLflow/Kubeflow/Vertex AI] still relevant in 2026?"
- "[MLOps tool A] vs [MLOps tool B]?"
- "Best way to deploy [LLM/ML model] to production?"
- "What feature store should I use?"
Minimal Recency Check
- Start from
data/sources.json and prefer sources with add_as_web_search: true.
- If web search or browsing is available, confirm at least: (a) the tool’s latest release/docs date, (b) active maintenance signals, (c) a recent comparison/alternatives post.
- If live search is not available, state that you are relying on static knowledge +
data/sources.json, and recommend validation steps (POC + evals + rollout plan).
What to Report
After searching, provide:
- Current landscape: What MLOps tools/platforms are popular NOW
- Emerging trends: New approaches gaining traction (LLMOps, GenAI ops)
- Deprecated/declining: Tools or approaches losing relevance
- Recommendation: Based on fresh data, not just static knowledge
Related Skills
For adjacent topics, reference these skills:
- ai-ml-data-science - EDA, feature engineering, modelling, evaluation, SQLMesh transformations
- ai-llm - Prompting, fine-tuning, evaluation for LLMs
- ai-agents - Agentic workflows, multi-agent systems, LLMOps
- ai-rag - RAG pipeline design, chunking, retrieval, evaluation
- ai-llm-inference - Model serving optimization, quantization, batching
- ai-prompt-engineering - Prompt design patterns and best practices
- data-lake-platform - Data lake/lakehouse infrastructure (ClickHouse, Iceberg, Kafka)
Use this skill to turn trained models into reliable services, not to derive the model itself.
1---2name: ai-mlops3description: Production MLOps and ML/LLM/agent security skill for deploying and operating ML systems in production (registry + CI/CD, serving, monitoring/drift, evaluation loops, incident response/runbooks, and governance), including GenAI security (prompt injection, jailbreaks, RAG security, privacy, and supply chain).4---5
6# MLOps & ML Security - Complete Reference (Jan 2026)
7
8Production ML lifecycle with **modern security practices**.
9
10This skill covers:
11
12- **Production**: Data ingestion, deployment, drift detection, monitoring, incident response
13- **Security**: Prompt injection, jailbreak defense, RAG security, output filtering
14- **Governance**: Privacy protection, supply chain security, safety evaluation
15
161. **Data ingestion** (dlt): Load data from APIs, databases to warehouses
172. **Model deployment**: Batch jobs, real-time APIs, hybrid systems, event-driven automation
183. **Operations**: Real-time monitoring, drift detection, automated retraining, incident response
19
20**Modern Best Practices (Jan 2026)**:
21
22- Version everything that can change: model artifacts, data snapshots, feature definitions, prompts/configs, and agent graphs; require reproducibility, rollbacks, and audit logs (NIST SSDF: https://csrc.nist.gov/pubs/sp/800/218/final).
23- Gate changes with evals (offline + online) and safe rollout (shadow/canary/blue-green); treat regressions in quality, safety, latency, and cost as release blockers.
24- Align controls and documentation to risk posture (EU AI Act: https://eur-lex.europa.eu/eli/reg/2024/1689/oj; NIST AI RMF + GenAI profile: https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf, https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf).
25- Operationalize security: threat model the full system (data, model, prompts, tools, RAG), harden the supply chain (SBOM/signing), and ship incident playbooks for both reliability and safety events.
26
27It is execution-focused:
28
29- Data ingestion patterns (REST APIs, database replication, incremental loading)
30- Deployment patterns (batch, online, hybrid, streaming, event-driven)
31- **Automated monitoring** with real-time drift detection
32- **Automated retraining** pipelines (monitor → detect → trigger → validate → deploy)
33- Incident handling with validated rollback and postmortems
34- Links to copy-paste templates in `assets/`
35
36## Quick Reference
37
38| Task | Tool/Framework | Command | When to Use |
39|------|----------------|---------|-------------|
40| Data Ingestion | dlt (data load tool) | `dlt pipeline run`, `dlt init` | Loading from APIs, databases to warehouses |
41| Batch Deployment | Airflow, Dagster, Prefect | `airflow dags trigger`, `dagster job launch` | Scheduled predictions on large datasets |
42| API Deployment | FastAPI, Flask, TorchServe | `uvicorn app:app`, `torchserve --start` | Real-time inference (<500ms latency) |
43| LLM Serving | vLLM, TGI, BentoML | `vllm serve model`, `bentoml serve` | High-throughput LLM inference |
44| Model Registry | MLflow, W&B, ZenML | `mlflow.register_model()`, `zenml model register` | Versioning and promoting models |
45| Drift Detection | Statistical tests + monitors | PSI/KS, embedding drift, prediction drift | Detect data/process changes and trigger review |
46| Monitoring | Prometheus, Grafana | `prometheus.yml`, Grafana dashboards | Metrics, alerts, SLO tracking |
47| AgentOps | AgentOps, Langfuse, LangSmith | `agentops.init()`, trace visualization | AI agent observability, session replay |
48| Incident Response | Runbooks, PagerDuty | Documented playbooks, alert routing | Handling failures and degradation |
49
50## Use This Skill When
51
52Use this skill when the user asks for **deployment, operations, monitoring, incident handling, or governance** for ML/LLM/agent systems, e.g.:
53
54- "How do I deploy this model to prod?"
55- "Design a batch + online scoring architecture."
56- "Add monitoring and drift detection to our model."
57- "Write an incident runbook for this ML service."
58- "Package this LLM/RAG pipeline as an API."
59- "Plan our retraining and promotion workflow."
60- "Load data from Stripe API to Snowflake."
61- "Set up incremental database replication with dlt."
62- "Build an ELT pipeline for warehouse loading."
63
64If the user is asking only about **EDA, modelling, or theory**, prefer:
65
66- `ai-ml-data-science` (EDA, features, modelling, SQL transformation with SQLMesh)
67- `ai-llm` (prompting, fine-tuning, eval)
68- `ai-rag` (retrieval pipeline design)
69- `ai-llm-inference` (compression, spec decode, serving internals)
70
71If the user is asking about **SQL transformation (after data is loaded)**, prefer:
72
73- `ai-ml-data-science` (SQLMesh templates for staging, intermediate, marts layers)
74
75## Decision Tree: Choosing Deployment Strategy
76
77```text
78User needs to deploy: [ML System]
79 ├─ Data Ingestion?
80 │ ├─ From REST APIs? → dlt REST API templates
81 │ ├─ From databases? → dlt database sources (PostgreSQL, MySQL, MongoDB)
82 │ └─ Incremental loading? → dlt incremental patterns (timestamp, ID-based)
83 │
84 ├─ Model Serving?
85 │ ├─ Latency <500ms? → FastAPI real-time API
86 │ ├─ Batch predictions? → Airflow/Dagster batch pipeline
87 │ └─ Mix of both? → Hybrid (batch features + online scoring)
88 │
89 ├─ Monitoring & Ops?
90 │ ├─ Drift detection? → Evidently + automated retraining triggers
91 │ ├─ Performance tracking? → Prometheus + Grafana dashboards
92 │ └─ Incident response? → Runbooks + PagerDuty alerts
93 │
94 └─ LLM/RAG Production?
95 ├─ Cost optimization? → Caching, prompt templates, token budgets
96 └─ Safety? → See ai-mlops skill
97```
98
99## Core Concepts (Vendor-Agnostic)
100
101- **Lifecycle loop**: train → validate → deploy → monitor → respond → retrain/retire.
102- **Risk controls**: access control, data minimization, logging, and change management (NIST AI RMF: https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf).
103- **Observability planes**: system metrics (latency/errors), data metrics (freshness/drift), quality metrics (model performance).
104- **Incident readiness**: detection, containment, rollback, and root-cause analysis.
105
106## Do / Avoid
107
108**Do**
109- Do gate deployments with repeatable checks: evaluation pass, load test, security review, rollback plan.
110- Do version everything: code, data, features, model artifact, prompt templates, configuration.
111- Do define SLOs and budgets (latency/cost/error rate) before optimizing.
112
113**Avoid**
114- Avoid manual “clickops” deployments without audit trail.
115- Avoid silent upgrades; require eval + canary for model/prompt changes.
116- Avoid drift dashboards without actions; every alert needs an owner and runbook.
117
118## Core Patterns Overview
119
120This skill provides production-ready patterns and guides organized into comprehensive references:
121
122### Data & Infrastructure Patterns
123
124**Pattern 0: Data Contracts, Ingestion & Lineage**
125→ See [Data Ingestion Patterns](references/data-ingestion-patterns.md)
126
127- Data contracts with SLAs and versioning
128- Ingestion modes (CDC, batch, streaming)
129- Lineage tracking and schema evolution
130- Replay and backfill procedures
131
132**Pattern 1: Choose Deployment Mode**
133→ See [Deployment Patterns](references/deployment-patterns.md)
134
135- Decision table (batch, online, hybrid, streaming)
136- When to use each mode
137- Deployment mode selection checklist
138
139**Pattern 2: Standard Deployment Lifecycle**
140→ See [Deployment Lifecycle](references/deployment-lifecycle.md)
141
142- Pre-deploy, deploy, observe, operate, evolve phases
143- Environment promotion (dev → staging → prod)
144- Gradual rollout strategies (canary, blue-green)
145
146**Pattern 3: Packaging & Model Registry**
147→ See [Model Registry Patterns](references/model-registry-patterns.md)
148
149- Model registry structure and metadata
150- Packaging strategies (Docker, ONNX, MLflow)
151- Promotion flows (experimental → production)
152- Versioning and governance
153
154### Serving Patterns
155
156**Pattern 4: Batch Scoring Pipeline**
157→ See [Deployment Patterns](references/deployment-patterns.md)
158
159- Orchestration with Airflow/Dagster
160- Idempotent scoring jobs
161- Validation and backfill procedures
162
163**Pattern 5: Real-Time API Scoring**
164→ See [API Design Patterns](references/api-design-patterns.md)
165
166- Service design (HTTP/JSON, gRPC)
167- Input/output schemas
168- Rate limiting, timeouts, circuit breakers
169
170**Pattern 6: Hybrid & Feature Store Integration**
171→ See [Feature Store Patterns](references/feature-store-patterns.md)
172
173- Batch vs online features
174- Feature store architecture
175- Training-serving consistency
176- Point-in-time correctness
177
178### Operations Patterns
179
180**Pattern 7: Monitoring & Alerting**
181→ See [Monitoring Best Practices](references/monitoring-best-practices.md)
182
183- Data, performance, and technical metrics
184- SLO definition and tracking
185- Dashboard design and alerting strategies
186
187**Pattern 8: Drift Detection & Automated Retraining**
188→ See [Drift Detection Guide](references/drift-detection-guide.md)
189
190- Automated retraining triggers
191- Event-driven retraining pipelines
192
193**Pattern 9: Incidents & Runbooks**
194→ See [Incident Response Playbooks](references/incident-response-playbooks.md)
195
196- Common failure modes
197- Detection, diagnosis, resolution
198- Post-mortem procedures
199
200**Pattern 10: LLM / RAG in Production**
201→ See [LLM & RAG Production Patterns](references/llm-rag-production-patterns.md)
202
203- Prompt and configuration management
204- Safety and compliance (PII, jailbreaks)
205- Cost optimization (token budgets, caching)
206- Monitoring and fallbacks
207
208**Pattern 11: Cross-Region, Residency & Rollback**
209→ See [Multi-Region Patterns](references/multi-region-patterns.md)
210
211- Multi-region deployment architectures
212- Data residency and tenant isolation
213- Disaster recovery and failover
214- Regional rollback procedures
215
216**Pattern 12: Online Evaluation & Feedback Loops**
217→ See [Online Evaluation Patterns](references/online-evaluation-patterns.md)
218
219- Feedback signal collection (implicit, explicit)
220- Shadow and canary deployments
221- A/B testing with statistical significance
222- Human-in-the-loop labeling
223- Automated retraining cadence
224
225**Pattern 13: AgentOps (AI Agent Operations)**
226→ See [AgentOps Patterns](references/agentops-patterns.md)
227
228- Session tracing and replay for AI agents
229- Cost and latency tracking across agent runs
230- Multi-agent visualization and debugging
231- Tool invocation monitoring
232- Integration with CrewAI, LangGraph, OpenAI Agents SDK
233
234**Pattern 14: Edge MLOps & TinyML**
235→ See [Edge MLOps Patterns](references/edge-mlops-patterns.md)
236
237- Device-aware CI/CD pipelines
238- OTA model updates with rollback
239- Federated learning operations
240- Edge drift detection
241- Intermittent connectivity handling
242
243## Resources (Detailed Guides)
244
245For comprehensive operational guides, see:
246
247**Core Infrastructure:**
248
249- **[Data Ingestion Patterns](references/data-ingestion-patterns.md)** - Data contracts, CDC, batch/streaming ingestion, lineage, schema evolution
250- **[Deployment Lifecycle](references/deployment-lifecycle.md)** - Pre-deploy validation, environment promotion, gradual rollout, rollback
251- **[Model Registry Patterns](references/model-registry-patterns.md)** - Versioning, packaging, promotion workflows, governance
252- **[Feature Store Patterns](references/feature-store-patterns.md)** - Batch/online features, hybrid architectures, consistency, latency optimization
253
254**Serving & APIs:**
255
256- **[Deployment Patterns](references/deployment-patterns.md)** - Batch, online, hybrid, streaming deployment strategies and architectures
257- **[API Design Patterns](references/api-design-patterns.md)** - ML/LLM/RAG API patterns, input/output schemas, reliability patterns, versioning
258
259**Operations & Reliability:**
260
261- **[Monitoring Best Practices](references/monitoring-best-practices.md)** - Metrics collection, alerting strategies, SLO definition, dashboard design
262- **[Drift Detection Guide](references/drift-detection-guide.md)** - Statistical tests, automated detection, retraining triggers, recovery strategies
263- **[Incident Response Playbooks](references/incident-response-playbooks.md)** - Runbooks for common failure modes, diagnostics, resolution steps
264
265**Security & Governance:**
266
267- **[Threat Models](references/threat-models.md)** - Trust boundaries, attack surface, control mapping
268- **[Prompt Injection Mitigation](references/prompt-injection-mitigation.md)** - Input hardening, tool/RAG containment, least privilege
269- **[Jailbreak Defense](references/jailbreak-defense.md)** - Robust refusal behavior, safe completion patterns
270- **[RAG Security](references/rag-security.md)** - Retrieval poisoning, context injection, sensitive data leakage
271- **[Output Filtering](references/output-filtering.md)** - Layered filters (PII/toxicity/policy), block/rewrite strategies
272- **[Privacy Protection](references/privacy-protection.md)** - PII handling, data minimization, retention, consent
273- **[Supply Chain Security](references/supply-chain-security.md)** - SBOM, dependency pinning, artifact signing
274- **[Safety Evaluation](references/safety-evaluation.md)** - Red teaming, eval sets, incident readiness
275
276**Advanced Patterns:**
277
278- **[LLM & RAG Production Patterns](references/llm-rag-production-patterns.md)** - Prompt management, safety, cost optimization, caching, monitoring
279- **[Multi-Region Patterns](references/multi-region-patterns.md)** - Multi-region deployment, data residency, disaster recovery, rollback
280- **[Online Evaluation Patterns](references/online-evaluation-patterns.md)** - A/B testing, shadow deployments, feedback loops, automated retraining
281- **[AgentOps Patterns](references/agentops-patterns.md)** - AI agent observability, session replay, cost tracking, multi-agent debugging
282- **[Edge MLOps Patterns](references/edge-mlops-patterns.md)** - TinyML, federated learning, OTA updates, device-aware CI/CD
283
284## Templates
285
286Use these as copy-paste starting points for production artifacts:
287
288### Data Ingestion (dlt)
289
290For loading data into warehouses and pipelines:
291
292- **[dlt basic pipeline setup](../data-lake-platform/assets/ingestion/dlt/template-dlt-pipeline.md)** - Install, configure, run basic extraction and loading
293- **[dlt REST API sources](../data-lake-platform/assets/ingestion/dlt/template-dlt-rest-api.md)** - Extract from REST APIs with pagination, authentication, rate limiting
294- **[dlt database sources](../data-lake-platform/assets/ingestion/dlt/template-dlt-database-source.md)** - Replicate from PostgreSQL, MySQL, MongoDB, SQL Server
295- **[dlt incremental loading](../data-lake-platform/assets/ingestion/dlt/template-dlt-incremental.md)** - Timestamp-based, ID-based, merge/upsert patterns, lookback windows
296- **[dlt warehouse loading](../data-lake-platform/assets/ingestion/dlt/template-dlt-warehouse-loading.md)** - Load to Snowflake, BigQuery, Redshift, Postgres, DuckDB
297
298**Use dlt when:**
299
300- Loading data from APIs (Stripe, HubSpot, Shopify, custom APIs)
301- Replicating databases to warehouses
302- Building ELT pipelines with incremental loading
303- Managing data ingestion with Python
304
305**For SQL transformation (after ingestion), use:**
306
307→ `ai-ml-data-science` skill (SQLMesh templates for staging/intermediate/marts layers)
308
309### Deployment & Packaging
310
311- **[Deployment & MLOps template](assets/deployment/template-deployment-mlops.md)** - Complete MLOps lifecycle, model registry, promotion workflows
312- **[Deployment readiness checklist](assets/deployment/deployment-readiness-checklist.md)** - Go/No-Go gate, monitoring, and rollback plan
313- **[API service template](assets/deployment/template-api-service.md)** - Real-time REST/gRPC API with FastAPI, input validation, rate limiting
314- **[Batch scoring pipeline template](assets/deployment/template-batch-pipeline.md)** - Orchestrated batch inference with Airflow/Dagster, validation, backfill
315
316### Monitoring & Operations
317
318- **[Monitoring & alerting template](assets/monitoring/template-monitoring-plan.md)** - Data/performance/technical metrics, dashboards, SLO definition
319- **[Drift detection & retraining template](assets/monitoring/template-drift-retraining.md)** - Automated drift detection, retraining triggers, promotion pipelines
320- **[Incident runbook template](assets/ops/template-incident-runbook.md)** - Failure mode playbooks, diagnosis steps, resolution procedures
321
322## Navigation
323
324**Resources**
325- [references/drift-detection-guide.md](references/drift-detection-guide.md)
326- [references/model-registry-patterns.md](references/model-registry-patterns.md)
327- [references/online-evaluation-patterns.md](references/online-evaluation-patterns.md)
328- [references/monitoring-best-practices.md](references/monitoring-best-practices.md)
329- [references/llm-rag-production-patterns.md](references/llm-rag-production-patterns.md)
330- [references/api-design-patterns.md](references/api-design-patterns.md)
331- [references/incident-response-playbooks.md](references/incident-response-playbooks.md)
332- [references/deployment-patterns.md](references/deployment-patterns.md)
333- [references/data-ingestion-patterns.md](references/data-ingestion-patterns.md)
334- [references/deployment-lifecycle.md](references/deployment-lifecycle.md)
335- [references/feature-store-patterns.md](references/feature-store-patterns.md)
336- [references/multi-region-patterns.md](references/multi-region-patterns.md)
337- [references/agentops-patterns.md](references/agentops-patterns.md)
338- [references/edge-mlops-patterns.md](references/edge-mlops-patterns.md)
339
340**Templates**
341- [template-dlt-pipeline.md](../data-lake-platform/assets/ingestion/dlt/template-dlt-pipeline.md)
342- [template-dlt-rest-api.md](../data-lake-platform/assets/ingestion/dlt/template-dlt-rest-api.md)
343- [template-dlt-database-source.md](../data-lake-platform/assets/ingestion/dlt/template-dlt-database-source.md)
344- [template-dlt-incremental.md](../data-lake-platform/assets/ingestion/dlt/template-dlt-incremental.md)
345- [template-dlt-warehouse-loading.md](../data-lake-platform/assets/ingestion/dlt/template-dlt-warehouse-loading.md)
346- [assets/deployment/template-deployment-mlops.md](assets/deployment/template-deployment-mlops.md)
347- [assets/deployment/deployment-readiness-checklist.md](assets/deployment/deployment-readiness-checklist.md)
348- [assets/deployment/template-api-service.md](assets/deployment/template-api-service.md)
349- [assets/deployment/template-batch-pipeline.md](assets/deployment/template-batch-pipeline.md)
350- [assets/ops/template-incident-runbook.md](assets/ops/template-incident-runbook.md)
351- [assets/monitoring/template-drift-retraining.md](assets/monitoring/template-drift-retraining.md)
352- [assets/monitoring/template-monitoring-plan.md](assets/monitoring/template-monitoring-plan.md)
353
354**Data**
355- [data/sources.json](data/sources.json) - Curated external references
356
357## External Resources
358
359See `data/sources.json` for curated references on:
360
361- Serving frameworks (FastAPI, Flask, gRPC, TorchServe, KServe, Ray Serve)
362- Orchestration (Airflow, Dagster, Prefect)
363- Model registries and MLOps (MLflow, W&B, Vertex AI, Sagemaker)
364- Monitoring and observability (Prometheus, Grafana, OpenTelemetry, Evidently)
365- Feature stores (Feast, Tecton, Vertex, Databricks)
366- Streaming & messaging (Kafka, Pulsar, Kinesis)
367- LLMOps & RAG infra (vector DBs, LLM gateways, safety tools)
368
369## Data Lake & Lakehouse
370
371For comprehensive data lake/lakehouse patterns (beyond dlt ingestion), see **[data-lake-platform](../data-lake-platform/SKILL.md)**:
372
373- **Table formats:** Apache Iceberg, Delta Lake, Apache Hudi
374- **Query engines:** ClickHouse, DuckDB, Apache Doris, StarRocks
375- **Alternative ingestion:** Airbyte (GUI-based connectors)
376- **Transformation:** dbt (alternative to SQLMesh)
377- **Streaming:** Apache Kafka patterns
378- **Orchestration:** Dagster, Airflow
379
380This skill focuses on **ML-specific deployment, monitoring, and security**. Use data-lake-platform for general-purpose data infrastructure.
381
382## Recency Protocol (Tooling Recommendations)
383
384When users ask recommendation questions about MLOps tooling, verify recency before answering.
385
386### Trigger Conditions
387
388- "What's the best MLOps platform for [use case]?"
389- "What should I use for [deployment/monitoring/drift detection]?"
390- "What's the latest in MLOps?"
391- "Current best practices for [model registry/feature store/observability]?"
392- "Is [MLflow/Kubeflow/Vertex AI] still relevant in 2026?"
393- "[MLOps tool A] vs [MLOps tool B]?"
394- "Best way to deploy [LLM/ML model] to production?"
395- "What feature store should I use?"
396
397### Minimal Recency Check
398
3991. Start from `data/sources.json` and prefer sources with `add_as_web_search: true`.
4002. If web search or browsing is available, confirm at least: (a) the tool’s latest release/docs date, (b) active maintenance signals, (c) a recent comparison/alternatives post.
4013. If live search is not available, state that you are relying on static knowledge + `data/sources.json`, and recommend validation steps (POC + evals + rollout plan).
402
403### What to Report
404
405After searching, provide:
406
407- **Current landscape**: What MLOps tools/platforms are popular NOW
408- **Emerging trends**: New approaches gaining traction (LLMOps, GenAI ops)
409- **Deprecated/declining**: Tools or approaches losing relevance
410- **Recommendation**: Based on fresh data, not just static knowledge
411
412## Related Skills
413
414For adjacent topics, reference these skills:
415
416- **[ai-ml-data-science](../ai-ml-data-science/SKILL.md)** - EDA, feature engineering, modelling, evaluation, SQLMesh transformations
417- **[ai-llm](../ai-llm/SKILL.md)** - Prompting, fine-tuning, evaluation for LLMs
418- **[ai-agents](../ai-agents/SKILL.md)** - Agentic workflows, multi-agent systems, LLMOps
419- **[ai-rag](../ai-rag/SKILL.md)** - RAG pipeline design, chunking, retrieval, evaluation
420- **[ai-llm-inference](../ai-llm-inference/SKILL.md)** - Model serving optimization, quantization, batching
421- **[ai-prompt-engineering](../ai-prompt-engineering/SKILL.md)** - Prompt design patterns and best practices
422- **[data-lake-platform](../data-lake-platform/SKILL.md)** - Data lake/lakehouse infrastructure (ClickHouse, Iceberg, Kafka)
423
424Use this skill to **turn trained models into reliable services**, not to derive the model itself.