LLM Development & Engineering — Complete Reference
Build, evaluate, and deploy LLM systems with modern production standards.
This skill covers the full LLM lifecycle:
- Development: Strategy selection, dataset design, instruction tuning, PEFT/LoRA fine-tuning
- Evaluation: Automated testing, LLM-as-judge, metrics, rollout gates
- Deployment: Serving handoff, latency/cost budgeting, reliability patterns (see
ai-llm-inference)
- Operations: Quality monitoring, change management, incident response (see
ai-mlops)
- Safety: Threat modeling, data governance, layered mitigations (NIST AI RMF: https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf)
Modern Best Practices (2026):
- Treat the model as a component with contracts, budgets, and rollback plans (not "magic").
- Separate core concepts (tokenization, context, training vs adaptation) from implementation choices (providers, SDKs).
- Gate upgrades with repeatable evals and staged rollout; avoid blind model swaps.
- Cost-aware engineering: Measure cost per successful outcome, not just cost per token; design tiering/caching early.
- Security-by-design: Threat model prompt injection, data leakage, and tool abuse; treat guardrails as production code.
For detailed patterns: See Resources and Templates sections below.
Quick Reference
| Task |
Tool/Framework |
Command/Pattern |
When to Use |
| Choose architecture |
Prompt vs RAG vs fine-tune |
Start simple; add retrieval/adaptation only if needed |
New products and migrations |
| Model selection |
Scoring matrix |
Quality/latency/cost/privacy/license weighting |
Provider changes and procurement |
| Cost optimization |
Tiered models + caching |
Cascade routing, prompt caching, budget guardrails |
Cost-sensitive production |
| Fine-tuning ROI |
ROI calculator |
Break-even analysis, TCO comparison |
Investment decisions |
| Prompt contracts |
Structured output + constraints |
JSON schema, max tokens, refusal rules |
Reliability and integration |
| RAG integration |
Hybrid retrieval + grounding |
Retrieve → rerank → pack → cite → verify |
Fresh/large corpora, traceability |
| Fine-tuning |
PEFT/LoRA (when justified) |
Small targeted datasets + regression suite |
Stable domains, repeated tasks |
| Evaluation |
Offline + online |
Golden sets + A/B + canary + monitoring |
Prevent regressions and drift |
Decision Tree: LLM System Architecture
Building LLM application: [Architecture Selection]
├─ Need current knowledge?
│ ├─ Simple Q&A? → Basic RAG (page-level chunking + hybrid retrieval)
│ └─ Complex retrieval? → Advanced RAG (reranking + contextual retrieval)
│
├─ Need tool use / actions?
│ ├─ Single task? → Simple agent (ReAct pattern)
│ └─ Multi-step workflow? → Multi-agent (LangGraph, CrewAI)
│
├─ Static behavior sufficient?
│ ├─ Quick MVP? → Prompt engineering (CI/CD integrated)
│ └─ Production quality? → Fine-tuning (PEFT/LoRA)
│
└─ Best results?
└─ Hybrid (RAG + Fine-tuning + Agents) → Comprehensive solution
See Decision Matrices for detailed selection criteria.
Cost-Quality Decision Framework
LLM spend is driven by usage-based inference (tokens/requests) plus supporting infra and engineering. Model selection is a cost-quality-latency-risk tradeoff.
Model Tier Strategy
| Tier | Typical profile | Use For |
|------|--------|------|---------|
| Value | Small/fast models | High-volume, simple tasks |
| Balanced | General-purpose models | Most production workloads |
| Premium | Frontier/large models | Hardest tasks, low volume |
Cost Optimization Levers
- Model tiering: Route simple requests to cheaper models (often large savings at scale)
- Prompt caching: Reuse stable prefixes/context (provider-specific discounts and constraints)
- Prompt optimization: Compress examples and instructions (typically meaningful token reduction)
- Output limits: Set appropriate max_tokens (prevents runaway costs)
When to Fine-Tune (ROI-Based)
Fine-tuning pays off when:
- Volume justifies it: >10k requests/month provides meaningful cost savings
- Domain is stable: Requirements unchanged for >6 months
- Data exists: >1,000 quality training examples available
- Break-even achievable: <12 months to recover investment
See Cost Economics for TCO modeling and Fine-Tuning ROI Calculator for investment analysis.
Core Concepts (Vendor-Agnostic)
- Model classes: encoder-only, decoder-only, encoder-decoder, multimodal; choose based on task and latency.
- Tokenization & limits: context window, max output, and prompt/template overhead drive both cost and tail latency.
- Adaptation options: prompting → retrieval → adapters (LoRA) → full fine-tune; choose by stability and ROI (LoRA: https://arxiv.org/abs/2106.09685).
- Evaluation: metrics must map to user value; report uncertainty and slice performance, not only global averages.
- Governance: data retention, residency, licensing, and auditability are product requirements (EU AI Act: https://eur-lex.europa.eu/eli/reg/2024/1689/oj; NIST GenAI Profile: https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf).
Implementation Practices (Tooling Examples)
- Use a provider abstraction (gateway/router) to enable fallbacks and staged upgrades.
- Instrument requests with tokens, latency, and error classes (OpenTelemetry GenAI semantic conventions: https://opentelemetry.io/docs/specs/semconv/gen-ai/).
- Maintain prompt/model registries with versioning, changelogs, and rollback criteria.
Do / Avoid
Do
- Do pin model + prompt versions in production, and re-run evals before any change.
- Do enforce budgets at the boundary: max tokens, max tools, max retries, max cost.
- Do plan for degraded modes (smaller model, cached answers, “unable to answer”).
Avoid
- Avoid model sprawl (unowned variants with no eval coverage).
- Avoid blind upgrades based on anecdotal quality; require measured impact.
- Avoid training on production logs without consent, governance, and leakage controls.
When to Use This Skill
Claude should invoke this skill when the user asks about:
- LLM preflight/project checklists, production best practices, or data pipelines
- Building or deploying RAG, agentic, or prompt-based LLM apps
- Prompt design, chain-of-thought (CoT), ReAct, or template patterns
- Troubleshooting LLM hallucination, bias, retrieval issues, or production failures
- Evaluating LLMs: benchmarks, multi-metric eval, or rollout/monitoring
- LLMOps: deployment, rollback, scaling, resource optimization
- Technology stack selection (models, vector DBs, frameworks)
- Production deployment strategies and operational patterns
Scope Boundaries (Use These Skills for Depth)
Resources (Best Practices & Operational Patterns)
Comprehensive operational guides with checklists, patterns, and decision frameworks:
Core Operational Patterns
Cost Economics & Decision Frameworks - Cost modeling, unit economics, TCO analysis
- Pricing/discount assumptions (verify against current provider docs)
- Cost-quality tradeoff framework and decision matrix
- Total Cost of Ownership (TCO) calculation
- Fine-tuning ROI framework and break-even analysis
- Prompt caching economics
- Cost monitoring and budget guardrails
Project Planning Patterns - Stack selection, FTI pipeline, performance budgeting
- AI engineering stack selection matrix
- Feature/Training/Inference (FTI) pipeline blueprint
- Performance budgeting and goodput gates
- Progressive complexity (prompt → RAG → fine-tune → hybrid)
Production Checklists - Pre-deployment validation and operational checklists
- LLM lifecycle checklist (modern production standards)
- Data & training, RAG pipeline, deployment & serving
- Safety/guardrails, evaluation, agentic systems
- Reliability & data infrastructure (DDIA-grade)
- Weekly production tasks
Common Design Patterns - Copy-paste ready implementation examples
- Chain-of-Thought (CoT) prompting
- ReAct (Reason + Act) pattern
- RAG pipeline (minimal to advanced)
- Agentic planning loop
- Self-reflection and multi-agent collaboration
Decision Matrices - Quick reference tables for selection
- RAG type decision matrix (naive → advanced → modular)
- Production evaluation table with targets and actions
- Model selection matrix (tier-based, vendor-agnostic)
- Vector database, embedding model, framework selection
- Deployment strategy matrix
Anti-Patterns - Common mistakes and prevention strategies
- Data leakage, prompt dilution, RAG context overload
- Agentic runaway, over-engineering, ignoring evaluation
- Hard-coded prompts, missing observability
- Detection methods and prevention code examples
Domain-Specific Patterns
- LLMOps Best Practices - Operational lifecycle and deployment patterns
- Evaluation Patterns - Testing, metrics, and quality validation
- Prompt Engineering Patterns - Quick reference (canonical skill: ai-prompt-engineering)
- Agentic Patterns - Quick reference (canonical skill: ai-agents)
- RAG Best Practices - Quick reference (canonical skill: ai-rag)
Note: Each resource file includes preflight/validation checklists, copy-paste reference tables, inline templates, anti-patterns, and decision matrices.
Templates (Copy-Paste Ready)
Production templates by use case and technology:
Selection & Governance
- Model Selection Matrix - Documented selection, scoring, licensing, and governance
- Fine-Tuning ROI Calculator - Investment analysis, break-even, go/no-go decisions
RAG Pipelines
- Basic RAG - Simple retrieval-augmented generation
- Advanced RAG - Hybrid retrieval, reranking, contextual embeddings
Prompt Engineering
- Chain-of-Thought - Step-by-step reasoning pattern
- ReAct - Reason + Act for tool use
Agentic Workflows
- Reflection Agent - Self-critique and improvement
- Multi-Agent - Manager-worker orchestration
Data Pipelines
- Data Quality - Validation, deduplication, PII detection
Deployment
- LLM Deployment - Production deployment with monitoring
Evaluation
- Multi-Metric Evaluation - Comprehensive testing suite
Shared Utilities (Centralized patterns — extract, don't duplicate)
Trend Awareness Protocol
IMPORTANT: For “best/latest” recommendations, verify recency using current sources (official docs/release notes/benchmarks). If you can’t browse, state assumptions and ask for timeframe + constraints.
Trigger Conditions
- "What's the best LLM model for [use case]?"
- "What should I use for [RAG/fine-tuning/agents]?"
- "What's the latest in LLM development?"
- "Current best practices for [prompting/evaluation/deployment]?"
- "Is [model/framework] still relevant in 2026?"
- "[Model A] vs [Model B]?" or "[Framework A] vs [Framework B]?"
- "Best vector database for [use case]?"
- "What agent framework should I use?"
Minimal Verification Checklist
- Confirm user constraints: latency, cost, privacy/compliance, deployment target, and toolchain.
- Check at least 2 authoritative sources from
data/sources.json (provider docs, release notes, pricing/quotas, deprecations).
- Prefer stable guidance (tradeoffs + decision criteria) over “one best model/framework”.
What to Report
After searching, provide:
- Current landscape: What models/frameworks are popular NOW (not 6 months ago)
- Emerging trends: New models, frameworks, or techniques gaining traction
- Deprecated/declining: Models/frameworks losing relevance or support
- Recommendation: Based on fresh data, not just static knowledge
Example Topics (verify with fresh sources)
- Latest frontier models (GPT-4.5, Claude 4, Gemini 2.x, Llama 4)
- Agent frameworks (LangGraph, CrewAI, AutoGen, Semantic Kernel)
- Vector databases (Pinecone, Qdrant, Weaviate, pgvector)
- RAG techniques (contextual retrieval, agentic RAG, graph RAG)
- Inference engines (vLLM, TensorRT-LLM, SGLang)
- Evaluation frameworks (RAGAS, DeepEval, Braintrust)
Related Skills
This skill integrates with complementary Claude Code skills:
Core Dependencies
- ai-rag - Retrieval pipelines: chunking, hybrid search, reranking, evaluation
- ai-prompt-engineering - Systematic prompt design, evaluation, testing, and optimization
- ai-agents - Agent architectures, tool use, multi-agent systems, autonomous workflows
Production & Operations
- ai-llm-inference - Production serving, quantization, batching, GPU optimization
- ai-mlops - Deployment, monitoring, incident response, security, and governance
External Resources
See data/sources.json for 50+ curated authoritative sources:
- Official LLM platform docs - OpenAI, Anthropic, Gemini, Mistral, Azure OpenAI, AWS Bedrock
- Open-source models and frameworks - HuggingFace Transformers, open-weight models, PEFT/LoRA, distributed training/inference stacks
- RAG frameworks and vector DBs - LlamaIndex, LangChain 1.2+, LangGraph, LangGraph Studio v2, Haystack, Pinecone, Qdrant, Chroma
- Agent frameworks (examples) - LangGraph, Semantic Kernel, AutoGen, CrewAI
- RAG innovations (examples) - Graph-based retrieval, hybrid retrieval, online evaluation loops
- Prompt engineering - Anthropic Prompt Library, Prompt Engineering Guide, CoT/ReAct patterns
- Evaluation and monitoring - OpenAI Evals, HELM, Anthropic Evals, LangSmith, W&B, Arize Phoenix
- Production deployment - Model gateways/routers, self-hosted serving, managed endpoints
Usage
For New Projects
- Start with Production Checklists - Validate all pre-deployment requirements
- Use Decision Matrices - Select technology stack
- Reference Project Planning Patterns - Design FTI pipeline
- Implement with Common Design Patterns - Copy-paste code examples
- Avoid Anti-Patterns - Learn from common mistakes
For Troubleshooting
- Check Anti-Patterns - Identify failure modes and mitigations
- Use Decision Matrices - Evaluate if architecture fits use case
- Reference Common Design Patterns - Verify implementation correctness
For Ongoing Operations
- Follow Production Checklists - Weekly operational tasks
- Integrate Evaluation Patterns - Continuous quality monitoring
- Apply LLMOps Best Practices - Deployment and rollback procedures
Navigation Summary
Quick Decisions: Decision Matrices
Pre-Deployment: Production Checklists
Planning: Project Planning Patterns
Implementation: Common Design Patterns
Troubleshooting: Anti-Patterns
Domain Depth: LLMOps | Evaluation | Prompts | Agents | RAG
Templates: assets/ - Copy-paste ready production code
Sources: data/sources.json - Authoritative documentation links
1---2name: ai-llm3description: Production LLM engineering skill. Covers strategy selection (prompting vs RAG vs fine-tuning), dataset design, PEFT/LoRA, evaluation workflows, deployment handoff to inference serving, and lifecycle operations with cost/safety controls.4---56# LLM Development & Engineering — Complete Reference78Build, evaluate, and deploy LLM systems with **modern production standards**.910This skill covers the full LLM lifecycle:1112- **Development**: Strategy selection, dataset design, instruction tuning, PEFT/LoRA fine-tuning13- **Evaluation**: Automated testing, LLM-as-judge, metrics, rollout gates14- **Deployment**: Serving handoff, latency/cost budgeting, reliability patterns (see `ai-llm-inference`)15- **Operations**: Quality monitoring, change management, incident response (see `ai-mlops`)16- **Safety**: Threat modeling, data governance, layered mitigations (NIST AI RMF: https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf)1718**Modern Best Practices (2026)**:1920- Treat the model as a **component** with contracts, budgets, and rollback plans (not "magic").21- Separate **core concepts** (tokenization, context, training vs adaptation) from **implementation choices** (providers, SDKs).22- Gate upgrades with repeatable evals and staged rollout; avoid blind model swaps.23- **Cost-aware engineering**: Measure cost per successful outcome, not just cost per token; design tiering/caching early.24- **Security-by-design**: Threat model prompt injection, data leakage, and tool abuse; treat guardrails as production code.2526**For detailed patterns:** See [Resources](#resources-best-practices--operational-patterns) and [Templates](#templates-copy-paste-ready) sections below.2728---2930## Quick Reference3132| Task | Tool/Framework | Command/Pattern | When to Use |33|------|----------------|-----------------|-------------|34| Choose architecture | Prompt vs RAG vs fine-tune | Start simple; add retrieval/adaptation only if needed | New products and migrations |35| Model selection | Scoring matrix | Quality/latency/cost/privacy/license weighting | Provider changes and procurement |36| **Cost optimization** | Tiered models + caching | Cascade routing, prompt caching, budget guardrails | Cost-sensitive production |37| **Fine-tuning ROI** | ROI calculator | Break-even analysis, TCO comparison | Investment decisions |38| Prompt contracts | Structured output + constraints | JSON schema, max tokens, refusal rules | Reliability and integration |39| RAG integration | Hybrid retrieval + grounding | Retrieve → rerank → pack → cite → verify | Fresh/large corpora, traceability |40| Fine-tuning | PEFT/LoRA (when justified) | Small targeted datasets + regression suite | Stable domains, repeated tasks |41| Evaluation | Offline + online | Golden sets + A/B + canary + monitoring | Prevent regressions and drift |4243---4445## Decision Tree: LLM System Architecture4647```text48Building LLM application: [Architecture Selection]49 ├─ Need current knowledge?50 │ ├─ Simple Q&A? → Basic RAG (page-level chunking + hybrid retrieval)51 │ └─ Complex retrieval? → Advanced RAG (reranking + contextual retrieval)52 │53 ├─ Need tool use / actions?54 │ ├─ Single task? → Simple agent (ReAct pattern)55 │ └─ Multi-step workflow? → Multi-agent (LangGraph, CrewAI)56 │57 ├─ Static behavior sufficient?58 │ ├─ Quick MVP? → Prompt engineering (CI/CD integrated)59 │ └─ Production quality? → Fine-tuning (PEFT/LoRA)60 │61 └─ Best results?62 └─ Hybrid (RAG + Fine-tuning + Agents) → Comprehensive solution63```6465**See [Decision Matrices](references/decision-matrices.md) for detailed selection criteria.**6667---6869## Cost-Quality Decision Framework7071LLM spend is driven by usage-based inference (tokens/requests) plus supporting infra and engineering. Model selection is a **cost-quality-latency-risk tradeoff**.7273### Model Tier Strategy7475| Tier | Typical profile | Use For |76|------|--------|------|---------|77| **Value** | Small/fast models | High-volume, simple tasks |78| **Balanced** | General-purpose models | Most production workloads |79| **Premium** | Frontier/large models | Hardest tasks, low volume |8081### Cost Optimization Levers82831. **Model tiering**: Route simple requests to cheaper models (often large savings at scale)842. **Prompt caching**: Reuse stable prefixes/context (provider-specific discounts and constraints)853. **Prompt optimization**: Compress examples and instructions (typically meaningful token reduction)864. **Output limits**: Set appropriate max_tokens (prevents runaway costs)8788### When to Fine-Tune (ROI-Based)8990Fine-tuning pays off when:91- **Volume justifies it**: >10k requests/month provides meaningful cost savings92- **Domain is stable**: Requirements unchanged for >6 months93- **Data exists**: >1,000 quality training examples available94- **Break-even achievable**: <12 months to recover investment9596**See [Cost Economics](references/cost-economics.md) for TCO modeling and [Fine-Tuning ROI Calculator](assets/selection/fine-tuning-roi-calculator.md) for investment analysis.**9798---99100## Core Concepts (Vendor-Agnostic)101102- **Model classes**: encoder-only, decoder-only, encoder-decoder, multimodal; choose based on task and latency.103- **Tokenization & limits**: context window, max output, and prompt/template overhead drive both cost and tail latency.104- **Adaptation options**: prompting → retrieval → adapters (LoRA) → full fine-tune; choose by stability and ROI (LoRA: https://arxiv.org/abs/2106.09685).105- **Evaluation**: metrics must map to user value; report uncertainty and slice performance, not only global averages.106- **Governance**: data retention, residency, licensing, and auditability are product requirements (EU AI Act: https://eur-lex.europa.eu/eli/reg/2024/1689/oj; NIST GenAI Profile: https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf).107108## Implementation Practices (Tooling Examples)109110- Use a **provider abstraction** (gateway/router) to enable fallbacks and staged upgrades.111- Instrument requests with tokens, latency, and error classes (OpenTelemetry GenAI semantic conventions: https://opentelemetry.io/docs/specs/semconv/gen-ai/).112- Maintain **prompt/model registries** with versioning, changelogs, and rollback criteria.113114## Do / Avoid115116**Do**117- Do pin model + prompt versions in production, and re-run evals before any change.118- Do enforce budgets at the boundary: max tokens, max tools, max retries, max cost.119- Do plan for degraded modes (smaller model, cached answers, “unable to answer”).120121**Avoid**122- Avoid model sprawl (unowned variants with no eval coverage).123- Avoid blind upgrades based on anecdotal quality; require measured impact.124- Avoid training on production logs without consent, governance, and leakage controls.125126## When to Use This Skill127128Claude should invoke this skill when the user asks about:129130- LLM preflight/project checklists, production best practices, or data pipelines131- Building or deploying RAG, agentic, or prompt-based LLM apps132- Prompt design, chain-of-thought (CoT), ReAct, or template patterns133- Troubleshooting LLM hallucination, bias, retrieval issues, or production failures134- Evaluating LLMs: benchmarks, multi-metric eval, or rollout/monitoring135- LLMOps: deployment, rollback, scaling, resource optimization136- Technology stack selection (models, vector DBs, frameworks)137- Production deployment strategies and operational patterns138139---140141## Scope Boundaries (Use These Skills for Depth)142143- **Prompt design & CI/CD** → [ai-prompt-engineering](../ai-prompt-engineering/SKILL.md)144- **RAG pipelines & chunking** → [ai-rag](../ai-rag/SKILL.md)145- **Search tuning (BM25, HNSW, hybrid)** → [ai-rag](../ai-rag/SKILL.md)146- **Agent architectures & tools** → [ai-agents](../ai-agents/SKILL.md)147- **Serving optimization/quantization** → [ai-llm-inference](../ai-llm-inference/SKILL.md)148- **Production deployment/monitoring** → [ai-mlops](../ai-mlops/SKILL.md)149- **Security/guardrails** → [ai-mlops](../ai-mlops/SKILL.md)150151---152153## Resources (Best Practices & Operational Patterns)154155Comprehensive operational guides with checklists, patterns, and decision frameworks:156157### Core Operational Patterns158159- **[Cost Economics & Decision Frameworks](references/cost-economics.md)** - Cost modeling, unit economics, TCO analysis160 - Pricing/discount assumptions (verify against current provider docs)161 - Cost-quality tradeoff framework and decision matrix162 - Total Cost of Ownership (TCO) calculation163 - Fine-tuning ROI framework and break-even analysis164 - Prompt caching economics165 - Cost monitoring and budget guardrails166167- **[Project Planning Patterns](references/project-planning-patterns.md)** - Stack selection, FTI pipeline, performance budgeting168 - AI engineering stack selection matrix169 - Feature/Training/Inference (FTI) pipeline blueprint170 - Performance budgeting and goodput gates171 - Progressive complexity (prompt → RAG → fine-tune → hybrid)172173- **[Production Checklists](references/production-checklists.md)** - Pre-deployment validation and operational checklists174 - LLM lifecycle checklist (modern production standards)175 - Data & training, RAG pipeline, deployment & serving176 - Safety/guardrails, evaluation, agentic systems177 - Reliability & data infrastructure (DDIA-grade)178 - Weekly production tasks179180- **[Common Design Patterns](references/common-design-patterns.md)** - Copy-paste ready implementation examples181 - Chain-of-Thought (CoT) prompting182 - ReAct (Reason + Act) pattern183 - RAG pipeline (minimal to advanced)184 - Agentic planning loop185 - Self-reflection and multi-agent collaboration186187- **[Decision Matrices](references/decision-matrices.md)** - Quick reference tables for selection188 - RAG type decision matrix (naive → advanced → modular)189 - Production evaluation table with targets and actions190 - Model selection matrix (tier-based, vendor-agnostic)191 - Vector database, embedding model, framework selection192 - Deployment strategy matrix193194- **[Anti-Patterns](references/anti-patterns.md)** - Common mistakes and prevention strategies195 - Data leakage, prompt dilution, RAG context overload196 - Agentic runaway, over-engineering, ignoring evaluation197 - Hard-coded prompts, missing observability198 - Detection methods and prevention code examples199200### Domain-Specific Patterns201202- **[LLMOps Best Practices](references/llmops-best-practices.md)** - Operational lifecycle and deployment patterns203- **[Evaluation Patterns](references/eval-patterns.md)** - Testing, metrics, and quality validation204- **[Prompt Engineering Patterns](references/prompt-engineering-patterns.md)** - Quick reference (canonical skill: [ai-prompt-engineering](../ai-prompt-engineering/SKILL.md))205- **[Agentic Patterns](references/agentic-patterns.md)** - Quick reference (canonical skill: [ai-agents](../ai-agents/SKILL.md))206- **[RAG Best Practices](references/rag-best-practices.md)** - Quick reference (canonical skill: [ai-rag](../ai-rag/SKILL.md))207208**Note:** Each resource file includes preflight/validation checklists, copy-paste reference tables, inline templates, anti-patterns, and decision matrices.209210---211212## Templates (Copy-Paste Ready)213214Production templates by use case and technology:215216### Selection & Governance217218- **[Model Selection Matrix](assets/selection/model-selection-matrix.md)** - Documented selection, scoring, licensing, and governance219- **[Fine-Tuning ROI Calculator](assets/selection/fine-tuning-roi-calculator.md)** - Investment analysis, break-even, go/no-go decisions220221### RAG Pipelines222223- **[Basic RAG](assets/rag-pipelines/template-basic-rag.md)** - Simple retrieval-augmented generation224- **[Advanced RAG](assets/rag-pipelines/template-advanced-rag.md)** - Hybrid retrieval, reranking, contextual embeddings225226### Prompt Engineering227228- **[Chain-of-Thought](assets/prompt-engineering/template-cot.md)** - Step-by-step reasoning pattern229- **[ReAct](assets/prompt-engineering/template-react.md)** - Reason + Act for tool use230231### Agentic Workflows232233- **[Reflection Agent](assets/agentic-workflows/template-reflection.md)** - Self-critique and improvement234- **[Multi-Agent](assets/agentic-workflows/template-multi-agent.md)** - Manager-worker orchestration235236### Data Pipelines237238- **[Data Quality](assets/data-pipelines/template-data-quality.md)** - Validation, deduplication, PII detection239240### Deployment241242- **[LLM Deployment](assets/deployment/template-llm-deployment.md)** - Production deployment with monitoring243244### Evaluation245246- **[Multi-Metric Evaluation](assets/evaluation/template-multi-metric.md)** - Comprehensive testing suite247248---249250## Shared Utilities (Centralized patterns — extract, don't duplicate)251252- [../software-clean-code-standard/utilities/llm-utilities.md](../software-clean-code-standard/utilities/llm-utilities.md) — Token counting, streaming, cost estimation253- [../software-clean-code-standard/utilities/error-handling.md](../software-clean-code-standard/utilities/error-handling.md) — Effect Result types, correlation IDs254- [../software-clean-code-standard/utilities/resilience-utilities.md](../software-clean-code-standard/utilities/resilience-utilities.md) — p-retry v6, circuit breaker for LLM API calls255- [../software-clean-code-standard/utilities/logging-utilities.md](../software-clean-code-standard/utilities/logging-utilities.md) — pino v9 + OpenTelemetry integration256- [../software-clean-code-standard/utilities/observability-utilities.md](../software-clean-code-standard/utilities/observability-utilities.md) — OpenTelemetry SDK, tracing, metrics257- [../software-clean-code-standard/utilities/config-validation.md](../software-clean-code-standard/utilities/config-validation.md) — Zod 3.24+, secrets management for API keys258- [../software-clean-code-standard/utilities/testing-utilities.md](../software-clean-code-standard/utilities/testing-utilities.md) — Test factories, fixtures, mocks259- [../software-clean-code-standard/references/clean-code-standard.md](../software-clean-code-standard/references/clean-code-standard.md) — Canonical clean code rules (`CC-*`) for citation260261---262263## Trend Awareness Protocol264265**IMPORTANT**: For “best/latest” recommendations, verify recency using current sources (official docs/release notes/benchmarks). If you can’t browse, state assumptions and ask for timeframe + constraints.266267### Trigger Conditions268269- "What's the best LLM model for [use case]?"270- "What should I use for [RAG/fine-tuning/agents]?"271- "What's the latest in LLM development?"272- "Current best practices for [prompting/evaluation/deployment]?"273- "Is [model/framework] still relevant in 2026?"274- "[Model A] vs [Model B]?" or "[Framework A] vs [Framework B]?"275- "Best vector database for [use case]?"276- "What agent framework should I use?"277278### Minimal Verification Checklist2792801. Confirm user constraints: latency, cost, privacy/compliance, deployment target, and toolchain.2812. Check at least 2 authoritative sources from `data/sources.json` (provider docs, release notes, pricing/quotas, deprecations).2823. Prefer stable guidance (tradeoffs + decision criteria) over “one best model/framework”.283284### What to Report285286After searching, provide:287288- **Current landscape**: What models/frameworks are popular NOW (not 6 months ago)289- **Emerging trends**: New models, frameworks, or techniques gaining traction290- **Deprecated/declining**: Models/frameworks losing relevance or support291- **Recommendation**: Based on fresh data, not just static knowledge292293### Example Topics (verify with fresh sources)294295- Latest frontier models (GPT-4.5, Claude 4, Gemini 2.x, Llama 4)296- Agent frameworks (LangGraph, CrewAI, AutoGen, Semantic Kernel)297- Vector databases (Pinecone, Qdrant, Weaviate, pgvector)298- RAG techniques (contextual retrieval, agentic RAG, graph RAG)299- Inference engines (vLLM, TensorRT-LLM, SGLang)300- Evaluation frameworks (RAGAS, DeepEval, Braintrust)301302---303304## Related Skills305306This skill integrates with complementary Claude Code skills:307308### Core Dependencies309310- **[ai-rag](../ai-rag/SKILL.md)** - Retrieval pipelines: chunking, hybrid search, reranking, evaluation311- **[ai-prompt-engineering](../ai-prompt-engineering/SKILL.md)** - Systematic prompt design, evaluation, testing, and optimization312- **[ai-agents](../ai-agents/SKILL.md)** - Agent architectures, tool use, multi-agent systems, autonomous workflows313314### Production & Operations315316- **[ai-llm-inference](../ai-llm-inference/SKILL.md)** - Production serving, quantization, batching, GPU optimization317- **[ai-mlops](../ai-mlops/SKILL.md)** - Deployment, monitoring, incident response, security, and governance318319---320321## External Resources322323See **[data/sources.json](data/sources.json)** for 50+ curated authoritative sources:324325- **Official LLM platform docs** - OpenAI, Anthropic, Gemini, Mistral, Azure OpenAI, AWS Bedrock326- **Open-source models and frameworks** - HuggingFace Transformers, open-weight models, PEFT/LoRA, distributed training/inference stacks327- **RAG frameworks and vector DBs** - LlamaIndex, LangChain 1.2+, LangGraph, LangGraph Studio v2, Haystack, Pinecone, Qdrant, Chroma328- **Agent frameworks (examples)** - LangGraph, Semantic Kernel, AutoGen, CrewAI329- **RAG innovations (examples)** - Graph-based retrieval, hybrid retrieval, online evaluation loops330- **Prompt engineering** - Anthropic Prompt Library, Prompt Engineering Guide, CoT/ReAct patterns331- **Evaluation and monitoring** - OpenAI Evals, HELM, Anthropic Evals, LangSmith, W&B, Arize Phoenix332- **Production deployment** - Model gateways/routers, self-hosted serving, managed endpoints333334---335336## Usage337338### For New Projects3393401. Start with **[Production Checklists](references/production-checklists.md)** - Validate all pre-deployment requirements3412. Use **[Decision Matrices](references/decision-matrices.md)** - Select technology stack3423. Reference **[Project Planning Patterns](references/project-planning-patterns.md)** - Design FTI pipeline3434. Implement with **[Common Design Patterns](references/common-design-patterns.md)** - Copy-paste code examples3445. Avoid **[Anti-Patterns](references/anti-patterns.md)** - Learn from common mistakes345346### For Troubleshooting3473481. Check **[Anti-Patterns](references/anti-patterns.md)** - Identify failure modes and mitigations3492. Use **[Decision Matrices](references/decision-matrices.md)** - Evaluate if architecture fits use case3503. Reference **[Common Design Patterns](references/common-design-patterns.md)** - Verify implementation correctness351352### For Ongoing Operations3533541. Follow **[Production Checklists](references/production-checklists.md)** - Weekly operational tasks3552. Integrate **[Evaluation Patterns](references/eval-patterns.md)** - Continuous quality monitoring3563. Apply **[LLMOps Best Practices](references/llmops-best-practices.md)** - Deployment and rollback procedures357358---359360## Navigation Summary361362**Quick Decisions:** [Decision Matrices](references/decision-matrices.md)363**Pre-Deployment:** [Production Checklists](references/production-checklists.md)364**Planning:** [Project Planning Patterns](references/project-planning-patterns.md)365**Implementation:** [Common Design Patterns](references/common-design-patterns.md)366**Troubleshooting:** [Anti-Patterns](references/anti-patterns.md)367368**Domain Depth:** [LLMOps](references/llmops-best-practices.md) | [Evaluation](references/eval-patterns.md) | [Prompts](references/prompt-engineering-patterns.md) | [Agents](references/agentic-patterns.md) | [RAG](references/rag-best-practices.md)369370**Templates:** [assets/](assets/) - Copy-paste ready production code371372**Sources:** [data/sources.json](data/sources.json) - Authoritative documentation links373374---