Rag Evaluation
Skill Profile
(Select at least one profile to enable specific modules)
Overview
Comprehensive guide to evaluating Retrieval-Augmented Generation systems, including retrieval metrics, generation quality, and end-to-end evaluation.
Why This Matters
- Quality Assurance: Ensures RAG systems produce accurate, relevant answers
- Performance Monitoring: Tracks retrieval and generation quality over time
- Optimization: Identifies areas for improvement in RAG pipelines
- Benchmarking: Enables comparison between different RAG implementations
- Trust: Builds confidence in RAG system outputs
Core Concepts & Rules
1. Core Principles
- Follow established patterns and conventions
- Maintain consistency across codebase
- Document decisions and trade-offs
2. Implementation Guidelines
- Start with the simplest viable solution
- Iterate based on feedback and requirements
- Test thoroughly before deployment
Inputs / Outputs / Contracts
- Inputs:
- Query dataset (questions with relevant documents and ground truth answers)
- RAG system components (retriever, generator, optional reranker)
- Evaluation configuration (K values, metrics to compute)
- Entry Conditions:
- RAG system is deployed and accessible
- Vector database is populated with documents
- Query dataset with ground truth is available
- LLM client for evaluation is configured
- Outputs:
- Retrieval metrics (precision@K, recall@K, MRR, NDCG, MAP)
- Generation metrics (faithfulness, relevance, coherence, factuality)
- End-to-end metrics (answer correctness, completeness)
- Evaluation report with visualizations
- Artifacts Required (Deliverables):
- Evaluation results JSON
- Metrics dashboard
- Comparison report (if comparing systems)
- Benchmark dataset
- Acceptance Evidence:
- All metrics computed successfully
- Results saved to file/database
- Dashboard displays metrics correctly
- Report generated with analysis
- Success Criteria:
- Retrieval precision@5 > 80%
- Generation faithfulness > 0.9
- End-to-end accuracy > 85%
- Evaluation completes within expected time
Skill Composition
- Depends on:
52-ai-evaluation/ground-truth-management
52-ai-evaluation/llm-judge-patterns
- Compatible with:
52-ai-evaluation/offline-vs-online-eval
52-ai-evaluation/regression-benchmarks
- Conflicts with: None
- Related Skills:
52-ai-evaluation/ground-truth-management
52-ai-evaluation/llm-judge-patterns
52-ai-evaluation/offline-vs-online-eval
52-ai-evaluation/regression-benchmarks
Quick Start / Implementation Example
- Review requirements and constraints
- Set up development environment
- Implement core functionality following patterns
- Write tests for critical paths
- Run tests and fix issues
- Document any deviations or decisions
# Example implementation following best practices
def example_function():
# Your implementation here
pass
Assumptions / Constraints / Non-goals
- Assumptions:
- Development environment is properly configured
- Required dependencies are available
- Team has basic understanding of domain
- Constraints:
- Must follow existing codebase conventions
- Time and resource limitations
- Compatibility requirements
- Non-goals:
- This skill does not cover edge cases outside scope
- Not a replacement for formal training
Compatibility & Prerequisites
- Supported Versions:
- Python 3.8+
- Node.js 16+
- Modern browsers (Chrome, Firefox, Safari, Edge)
- Required AI Tools:
- Code editor (VS Code recommended)
- Testing framework appropriate for language
- Version control (Git)
- Dependencies:
- Language-specific package manager
- Build tools
- Testing libraries
- Environment Setup:
.env.example keys: API_KEY, DATABASE_URL (no values)
Test Scenario Matrix (QA Strategy)
| Type |
Focus Area |
Required Scenarios / Mocks |
| Unit |
Core Logic |
Must cover primary logic and at least 3 edge/error cases. Target minimum 80% coverage |
| Integration |
DB / API |
All external API calls or database connections must be mocked during unit tests |
| E2E |
User Journey |
Critical user flows to test |
| Performance |
Latency / Load |
Benchmark requirements |
| Security |
Vuln / Auth |
SAST/DAST or dependency audit |
| Frontend |
UX / A11y |
Accessibility checklist (WCAG), Performance Budget (Lighthouse score) |
Technical Guardrails & Security Threat Model
1. Security & Privacy (Threat Model)
- Top Threats: Injection attacks, authentication bypass, data exposure
2. Performance & Resources
3. Architecture & Scalability
4. Observability & Reliability
Agent Directives & Error Recovery
(ข้อกำหนดสำหรับ AI Agent ในการคิดและแก้ปัญหาเมื่อเกิดข้อผิดพลาด)
- Thinking Process: Analyze root cause before fixing. Do not brute-force.
- Fallback Strategy: Stop after 3 failed test attempts. Output root cause and ask for human intervention/clarification.
- Self-Review: Check against Guardrails & Anti-patterns before finalizing.
- Output Constraints: Output ONLY the modified code block. Do not explain unless asked.
Definition of Done (DoD) Checklist
Anti-patterns / Pitfalls
- ⛔ Don't: Log PII, catch-all exception, N+1 queries
- ⚠️ Watch out for: Common symptoms and quick fixes
- 💡 Instead: Use proper error handling, pagination, and logging
Reference Links & Examples
- Internal documentation and examples
- Official documentation and best practices
- Community resources and discussions
Versioning & Changelog
- Version: 1.0.0
- Changelog:
- 2026-02-22: Initial version with complete template structure
Converted and distributed by TomeVault — claim your Tome and manage your conversions.
1---2name: rag-evaluation3description: Comprehensive guide to evaluating Retrieval-Augmented Generation systems, Use when this capability is needed.4---56# Rag Evaluation78## Skill Profile9*(Select at least one profile to enable specific modules)*10- [ ] **DevOps**11- [x] **Backend**12- [ ] **Frontend**13- [ ] **AI-RAG**14- [ ] **Security Critical**1516## Overview17Comprehensive guide to evaluating Retrieval-Augmented Generation systems, including retrieval metrics, generation quality, and end-to-end evaluation.1819## Why This Matters20- **Quality Assurance**: Ensures RAG systems produce accurate, relevant answers21- **Performance Monitoring**: Tracks retrieval and generation quality over time22- **Optimization**: Identifies areas for improvement in RAG pipelines23- **Benchmarking**: Enables comparison between different RAG implementations24- **Trust**: Builds confidence in RAG system outputs2526---2728## Core Concepts & Rules2930### 1. Core Principles31- Follow established patterns and conventions32- Maintain consistency across codebase33- Document decisions and trade-offs3435### 2. Implementation Guidelines36- Start with the simplest viable solution37- Iterate based on feedback and requirements38- Test thoroughly before deployment394041## Inputs / Outputs / Contracts42* **Inputs**:43 - Query dataset (questions with relevant documents and ground truth answers)44 - RAG system components (retriever, generator, optional reranker)45 - Evaluation configuration (K values, metrics to compute)46* **Entry Conditions**:47 - RAG system is deployed and accessible48 - Vector database is populated with documents49 - Query dataset with ground truth is available50 - LLM client for evaluation is configured51* **Outputs**:52 - Retrieval metrics (precision@K, recall@K, MRR, NDCG, MAP)53 - Generation metrics (faithfulness, relevance, coherence, factuality)54 - End-to-end metrics (answer correctness, completeness)55 - Evaluation report with visualizations56* **Artifacts Required (Deliverables)**:57 - Evaluation results JSON58 - Metrics dashboard59 - Comparison report (if comparing systems)60 - Benchmark dataset61* **Acceptance Evidence**:62 - All metrics computed successfully63 - Results saved to file/database64 - Dashboard displays metrics correctly65 - Report generated with analysis66* **Success Criteria**:67 - Retrieval precision@5 > 80%68 - Generation faithfulness > 0.969 - End-to-end accuracy > 85%70 - Evaluation completes within expected time7172## Skill Composition73* **Depends on**: 74 - [`52-ai-evaluation/ground-truth-management`](52-ai-evaluation/ground-truth-management/SKILL.md)75 - [`52-ai-evaluation/llm-judge-patterns`](52-ai-evaluation/llm-judge-patterns/SKILL.md)76* **Compatible with**: 77 - [`52-ai-evaluation/offline-vs-online-eval`](52-ai-evaluation/offline-vs-online-eval/SKILL.md)78 - [`52-ai-evaluation/regression-benchmarks`](52-ai-evaluation/regression-benchmarks/SKILL.md)79* **Conflicts with**: None80* **Related Skills**: 81 - [`52-ai-evaluation/ground-truth-management`](52-ai-evaluation/ground-truth-management/SKILL.md)82 - [`52-ai-evaluation/llm-judge-patterns`](52-ai-evaluation/llm-judge-patterns/SKILL.md)83 - [`52-ai-evaluation/offline-vs-online-eval`](52-ai-evaluation/offline-vs-online-eval/SKILL.md)84 - [`52-ai-evaluation/regression-benchmarks`](52-ai-evaluation/regression-benchmarks/SKILL.md)8586---8788## Quick Start / Implementation Example89901. Review requirements and constraints912. Set up development environment923. Implement core functionality following patterns934. Write tests for critical paths945. Run tests and fix issues956. Document any deviations or decisions9697```python98# Example implementation following best practices99def example_function():100 # Your implementation here101 pass102```103104105## Assumptions / Constraints / Non-goals106107* **Assumptions**:108 - Development environment is properly configured109 - Required dependencies are available110 - Team has basic understanding of domain111* **Constraints**:112 - Must follow existing codebase conventions113 - Time and resource limitations114 - Compatibility requirements115* **Non-goals**:116 - This skill does not cover edge cases outside scope117 - Not a replacement for formal training118119120## Compatibility & Prerequisites121122* **Supported Versions**:123 - Python 3.8+124 - Node.js 16+125 - Modern browsers (Chrome, Firefox, Safari, Edge)126* **Required AI Tools**:127 - Code editor (VS Code recommended)128 - Testing framework appropriate for language129 - Version control (Git)130* **Dependencies**:131 - Language-specific package manager132 - Build tools133 - Testing libraries134* **Environment Setup**:135 - `.env.example` keys: `API_KEY`, `DATABASE_URL` (no values)136137138## Test Scenario Matrix (QA Strategy)139140| Type | Focus Area | Required Scenarios / Mocks |141| :--- | :--- | :--- |142| **Unit** | Core Logic | Must cover primary logic and at least 3 edge/error cases. Target minimum 80% coverage |143| **Integration** | DB / API | All external API calls or database connections must be mocked during unit tests |144| **E2E** | User Journey | Critical user flows to test |145| **Performance** | Latency / Load | Benchmark requirements |146| **Security** | Vuln / Auth | SAST/DAST or dependency audit |147| **Frontend** | UX / A11y | Accessibility checklist (WCAG), Performance Budget (Lighthouse score) |148149150## Technical Guardrails & Security Threat Model151152### 1. Security & Privacy (Threat Model)153* **Top Threats**: Injection attacks, authentication bypass, data exposure154- [ ] **Data Handling**: Sanitize all user inputs to prevent Injection attacks. Never log raw PII155- [ ] **Secrets Management**: No hardcoded API keys. Use Env Vars/Secrets Manager156- [ ] **Authorization**: Validate user permissions before state changes157158### 2. Performance & Resources159- [ ] **Execution Efficiency**: Consider time complexity for algorithms160- [ ] **Memory Management**: Use streams/pagination for large data161- [ ] **Resource Cleanup**: Close DB connections/file handlers in finally blocks162163### 3. Architecture & Scalability164- [ ] **Design Pattern**: Follow SOLID principles, use Dependency Injection165- [ ] **Modularity**: Decouple logic from UI/Frameworks166167### 4. Observability & Reliability168- [ ] **Logging Standards**: Structured JSON, include trace IDs `request_id`169- [ ] **Metrics**: Track `error_rate`, `latency`, `queue_depth`170- [ ] **Error Handling**: Standardized error codes, no bare except171- [ ] **Observability Artifacts**:172 - **Log Fields**: timestamp, level, message, request_id173 - **Metrics**: request_count, error_count, response_time174 - **Dashboards/Alerts**: High Error Rate > 5%175176177## Agent Directives & Error Recovery178*(ข้อกำหนดสำหรับ AI Agent ในการคิดและแก้ปัญหาเมื่อเกิดข้อผิดพลาด)*179180- **Thinking Process**: Analyze root cause before fixing. Do not brute-force.181- **Fallback Strategy**: Stop after 3 failed test attempts. Output root cause and ask for human intervention/clarification.182- **Self-Review**: Check against Guardrails & Anti-patterns before finalizing.183- **Output Constraints**: Output ONLY the modified code block. Do not explain unless asked.184185186## Definition of Done (DoD) Checklist187188- [ ] Tests passed + coverage met189- [ ] Lint/Typecheck passed190- [ ] Logging/Metrics/Trace implemented191- [ ] Security checks passed192- [ ] Documentation/Changelog updated193- [ ] Accessibility/Performance requirements met (if frontend)194195196## Anti-patterns / Pitfalls197198* ⛔ **Don't**: Log PII, catch-all exception, N+1 queries199* ⚠️ **Watch out for**: Common symptoms and quick fixes200* 💡 **Instead**: Use proper error handling, pagination, and logging201202203## Reference Links & Examples204205* Internal documentation and examples206* Official documentation and best practices207* Community resources and discussions208209210## Versioning & Changelog211212* **Version**: 1.0.0213* **Changelog**:214 - 2026-02-22: Initial version with complete template structure215216---217> Converted and distributed by [TomeVault](https://tomevault.io/claim/amnadtaowsoam) — claim your Tome and manage your conversions.218<!-- tomevault:4.0:skill_md:2026-04-13 -->