RAG Patterns — Retrieval-Augmented Generation
Operating contract
Inputs
| Input |
Required |
Purpose |
| Domain evidence |
yes |
question set, authoritative corpus, tenancy and access rules, freshness target, citation needs, and evaluation set |
Outputs
- Produce: retrieval architecture, ingestion/chunking policy, index and filter contract, grounded-answer schema, and evaluation report.
Capability and permission boundaries
Default to read-only analysis. Read only scoped records; redact secrets and regulated data. Writes, execution, network calls, production configuration, customer communication, billing changes, and delegation require explicit authority and an identified owner. Never widen tenant, time-window, or system scope implicitly.
Degraded mode
When required telemetry, evidence, execution, network access, or write authority is unavailable, return a partial result with each unassessed item labelled, preserve the safest existing state, and state the evidence or approval needed to continue. Never convert missing evidence into a pass.
Decision rules
| Condition |
Action |
| Scope, owner, or threshold is missing |
Stop the affected decision and request it |
| Evidence is incomplete but read-only analysis is safe |
Produce a qualified partial result and gap list |
| A mutation exceeds authority or tenant boundary |
Block it and route for approval |
| Evidence meets the stated threshold |
Issue the output with provenance and owner |
Anti-Patterns
- Treating absent evidence as success. Fix: mark the check unassessed and name the missing source.
- Expanding one tenant or workflow to all tenants. Fix: enforce supplied scope at every query and action.
- Performing a production write during analysis. Fix: emit a reviewed change plan until authority is explicit.
- Reporting a metric without population, window, or source. Fix: attach all three.
- Hiding a failed threshold inside an average. Fix: report failure slices and the remediation owner.
Acknowledgement: Shared by Peter Bamuhigire, techguypeter.com, +256 784 464178.
Use When
- Use when building features that answer questions from private data, documents, policies, or time-sensitive information — RAG architecture, chunking strategies, hybrid search, re-ranking, vector databases, evaluation, agentic RAG, multimodal RAG...
Evidence Produced
| Category |
Artifact |
Format |
Example |
| Correctness |
RAG retrieval evaluation report |
Markdown doc covering recall / precision / answer-quality on a fixed eval set |
docs/ai/rag-eval-2026-04-16.md |
| Data safety |
Index ingestion + tenancy isolation note |
Markdown doc covering chunking, source filtering, and per-tenant index segregation |
docs/ai/rag-tenancy-note.md |
References
- Use the
references/ directory for deep detail after reading the core workflow below.
Overview
RAG solves the core LLM limitation: they only know what they were trained on. Use RAG to inject private data (invoices, menus, policies, reports) into every AI response.
Core principle: RAG = look up a database + LLM synthesises the results. The LLM never needs to "know" your data.
When to Use RAG
| Condition |
Action |
| Knowledge base < 200K tokens (~500 pages) |
Include everything in context — no RAG needed |
| Knowledge base > 200K tokens |
Use RAG |
| Data changes frequently (menus, prices, stock) |
RAG (update documents, not model) |
| Data is private/confidential |
RAG (keeps data out of training pipelines) |
| Need source citations |
RAG (chunks are traceable to source) |
| Model needs brand voice / domain jargon |
Fine-tune instead |
RAG vs Fine-Tuning
| Factor |
RAG |
Fine-Tuning |
| Up-to-date content |
✅ Yes (add docs anytime) |
❌ Stale until retrained |
| Hallucinations |
✅ Lower (document-grounded) |
❌ Higher |
| Source citations |
✅ Yes |
❌ No |
| Brand voice control |
❌ Weak |
✅ Strong |
| Domain jargon |
❌ Weak |
✅ Strong |
| Up-front cost |
✅ Lower |
❌ High |
Default: start with RAG. Fine-tune only when RAG + prompt engineering cannot deliver the required tone or vocabulary.
Additional Guidance
Guidance is split across two reference files so this entrypoint stays compact.
references/skill-deep-dive.md — architecture, chunking, retrieval, schema:
Pipeline Architecture
Chunking Strategies
Embedding Model Selection
Vector Database Selection
Retrieval Algorithms
Re-Ranking
Full RAG Query Algorithm
Query Rewriting (Multi-Turn)
RAG Schema (Multi-Tenant)
Evaluation Framework
Production Patterns
Agentic RAG
Multimodal RAG, Edge Cases, Cost Optimisation, Sources
references/production-rag.md — the progression from draft to production and the gates before shipping:
RAG Maturity Model — Naive → Advanced → Modular
Query Transformation — HyDE, Multi-Query, Step-Back
Contextual Compression
Self-RAG
RAGAS Evaluation — 4 metrics with production thresholds
Embedding Pipeline — batching, upserts, re-embed triggers, $/1M-token table
Cost Management Decision Tree — concrete dollar figures per branch
Failure Mode Playbook — empty, irrelevant, hallucinated, stale
Gates Before Shipping
Load the production file when building a RAG system that has to pass evaluation gates, survive multi-tenant review, or hit a cost budget under load.
Multi-Tenant Addendum
This skill describes RAG patterns in general. When the RAG feature ships inside a multi-tenant SaaS, the production answer is ai-rag-multi-tenant — per-tenant ingestion pipelines, vector store partitioning, tier-specific chunking and embedding models, defence-in-depth retrieval security, and citation grounding tied to live sources.
Cross-references:
ai-rag-multi-tenant — multi-tenant RAG end-to-end.
ai-tenant-isolation-patterns — vector-store partitioning tradeoffs and data-bleed tests.
ai-on-saas-architecture — KB service as a control-plane service.
ai-hallucination-slo-and-grounding — citation grounding + faithfulness SLO.
ai-model-gateway — gateway-mediated retrieval calls.
saas-tenant-data-portability-and-erasure — KB erasure cascade for embeddings.
Consolidated Child References
- Load references/routing.md to map retired AI child skill slugs to their reference modules.
1---2name: ai-rag-patterns3description: Use when building features that answer questions from private data, documents, policies, or time-sensitive information — RAG architecture, chunking strategies, hybrid search, re-ranking, vector databases, evaluation, agentic RAG, multimodal RAG...4---56# RAG Patterns — Retrieval-Augmented Generation78## Operating contract910## Inputs1112| Input | Required | Purpose |13|---|---|---|14| Domain evidence | yes | question set, authoritative corpus, tenancy and access rules, freshness target, citation needs, and evaluation set |1516## Outputs1718- Produce: retrieval architecture, ingestion/chunking policy, index and filter contract, grounded-answer schema, and evaluation report.1920## Capability and permission boundaries2122Default to read-only analysis. Read only scoped records; redact secrets and regulated data. Writes, execution, network calls, production configuration, customer communication, billing changes, and delegation require explicit authority and an identified owner. Never widen tenant, time-window, or system scope implicitly.2324## Degraded mode2526When required telemetry, evidence, execution, network access, or write authority is unavailable, return a partial result with each unassessed item labelled, preserve the safest existing state, and state the evidence or approval needed to continue. Never convert missing evidence into a pass.2728## Decision rules2930| Condition | Action |31|---|---|32| Scope, owner, or threshold is missing | Stop the affected decision and request it |33| Evidence is incomplete but read-only analysis is safe | Produce a qualified partial result and gap list |34| A mutation exceeds authority or tenant boundary | Block it and route for approval |35| Evidence meets the stated threshold | Issue the output with provenance and owner |3637## Anti-Patterns3839- Treating absent evidence as success. Fix: mark the check unassessed and name the missing source.40- Expanding one tenant or workflow to all tenants. Fix: enforce supplied scope at every query and action.41- Performing a production write during analysis. Fix: emit a reviewed change plan until authority is explicit.42- Reporting a metric without population, window, or source. Fix: attach all three.43- Hiding a failed threshold inside an average. Fix: report failure slices and the remediation owner.4445Acknowledgement: Shared by Peter Bamuhigire, techguypeter.com, +256 784 464178.4647<!-- dual-compat-start -->48## Use When4950- Use when building features that answer questions from private data, documents, policies, or time-sensitive information — RAG architecture, chunking strategies, hybrid search, re-ranking, vector databases, evaluation, agentic RAG, multimodal RAG...5152## Evidence Produced5354| Category | Artifact | Format | Example |55|----------|----------|--------|---------|56| Correctness | RAG retrieval evaluation report | Markdown doc covering recall / precision / answer-quality on a fixed eval set | `docs/ai/rag-eval-2026-04-16.md` |57| Data safety | Index ingestion + tenancy isolation note | Markdown doc covering chunking, source filtering, and per-tenant index segregation | `docs/ai/rag-tenancy-note.md` |5859## References6061- Use the `references/` directory for deep detail after reading the core workflow below.62<!-- dual-compat-end -->63## Overview6465RAG solves the core LLM limitation: they only know what they were trained on. Use RAG to inject private data (invoices, menus, policies, reports) into every AI response.6667**Core principle:** RAG = look up a database + LLM synthesises the results. The LLM never needs to "know" your data.6869---7071## When to Use RAG7273| Condition | Action |74|---|---|75| Knowledge base < 200K tokens (~500 pages) | Include everything in context — no RAG needed |76| Knowledge base > 200K tokens | Use RAG |77| Data changes frequently (menus, prices, stock) | RAG (update documents, not model) |78| Data is private/confidential | RAG (keeps data out of training pipelines) |79| Need source citations | RAG (chunks are traceable to source) |80| Model needs brand voice / domain jargon | Fine-tune instead |8182---8384## RAG vs Fine-Tuning8586| Factor | RAG | Fine-Tuning |87|---|---|---|88| Up-to-date content | ✅ Yes (add docs anytime) | ❌ Stale until retrained |89| Hallucinations | ✅ Lower (document-grounded) | ❌ Higher |90| Source citations | ✅ Yes | ❌ No |91| Brand voice control | ❌ Weak | ✅ Strong |92| Domain jargon | ❌ Weak | ✅ Strong |93| Up-front cost | ✅ Lower | ❌ High |9495**Default: start with RAG.** Fine-tune only when RAG + prompt engineering cannot deliver the required tone or vocabulary.9697---9899## Additional Guidance100101Guidance is split across two reference files so this entrypoint stays compact.102103**[references/skill-deep-dive.md](references/skill-deep-dive.md)** — architecture, chunking, retrieval, schema:104105- `Pipeline Architecture`106- `Chunking Strategies`107- `Embedding Model Selection`108- `Vector Database Selection`109- `Retrieval Algorithms`110- `Re-Ranking`111- `Full RAG Query Algorithm`112- `Query Rewriting (Multi-Turn)`113- `RAG Schema (Multi-Tenant)`114- `Evaluation Framework`115- `Production Patterns`116- `Agentic RAG`117- `Multimodal RAG`, `Edge Cases`, `Cost Optimisation`, `Sources`118119**[references/production-rag.md](references/production-rag.md)** — the progression from draft to production and the gates before shipping:120121- `RAG Maturity Model` — Naive → Advanced → Modular122- `Query Transformation` — HyDE, Multi-Query, Step-Back123- `Contextual Compression`124- `Self-RAG`125- `RAGAS Evaluation` — 4 metrics with production thresholds126- `Embedding Pipeline` — batching, upserts, re-embed triggers, $/1M-token table127- `Cost Management Decision Tree` — concrete dollar figures per branch128- `Failure Mode Playbook` — empty, irrelevant, hallucinated, stale129- `Gates Before Shipping`130131Load the production file when building a RAG system that has to pass evaluation gates, survive multi-tenant review, or hit a cost budget under load.132## Multi-Tenant Addendum133134This skill describes RAG patterns in general. When the RAG feature ships inside a multi-tenant SaaS, the production answer is `ai-rag-multi-tenant` — per-tenant ingestion pipelines, vector store partitioning, tier-specific chunking and embedding models, defence-in-depth retrieval security, and citation grounding tied to live sources.135136Cross-references:137- `ai-rag-multi-tenant` — multi-tenant RAG end-to-end.138- `ai-tenant-isolation-patterns` — vector-store partitioning tradeoffs and data-bleed tests.139- `ai-on-saas-architecture` — KB service as a control-plane service.140- `ai-hallucination-slo-and-grounding` — citation grounding + faithfulness SLO.141- `ai-model-gateway` — gateway-mediated retrieval calls.142- `saas-tenant-data-portability-and-erasure` — KB erasure cascade for embeddings.143## Consolidated Child References144145- Load [references/routing.md](references/routing.md) to map retired AI child skill slugs to their reference modules.