Ab Testing Ml
Skill Profile
(Select at least one profile to enable specific modules)
Overview
A/B testing compares ML models in production. This guide covers experiment design, statistical testing, multi-armed bandits, and monitoring for safely deploying and comparing ML models to determine which performs better on business metrics.
Why This Matters
- Risk Mitigation: A/B testing allows safe model deployment with gradual rollout
- Data-Driven Decisions: Statistical testing ensures decisions are based on significant results
- Business Impact: Multi-armed bandits optimize for business metrics during experimentation
Core Concepts & Rules
1. Core Principles
- Follow established patterns and conventions
- Maintain consistency across codebase
- Document decisions and trade-offs
2. Implementation Guidelines
- Start with the simplest viable solution
- Iterate based on feedback and requirements
- Test thoroughly before deployment
Inputs / Outputs / Contracts
- Inputs:
- User ID for consistent variant assignment
- Features for model prediction
- Actual outcomes for metric calculation
- Entry Conditions:
- Two or more models ready for production
- Sufficient traffic for statistical significance
- Metrics tracking infrastructure in place
- Outputs:
- Variant assignment (A/B)
- Model prediction with variant info
- Statistical test results (p-value, significance)
- Winner recommendation
- Artifacts Required (Deliverables):
- Experiment design configuration
- Traffic splitting logic
- Statistical testing implementation
- Multi-armed bandit algorithms
- Monitoring dashboard
- Acceptance Evidence:
- A/B test API endpoints functional
- Statistical tests passing
- Monitoring dashboard showing results
- Success Criteria:
- Statistical significance (p < 0.05)
- Minimum sample size achieved
- Business metrics improvement validated
Skill Composition
- Depends on: ML models, Metrics tracking, Feature flags
- Compatible with: ML Serving, Model Experiments, Feature Engineering
- Conflicts with: None
- Related Skills: ml-serving, model-experiments, feature-engineering
Quick Start / Implementation Example
- Review requirements and constraints
- Set up development environment
- Implement core functionality following patterns
- Write tests for critical paths
- Run tests and fix issues
- Document any deviations or decisions
# Example implementation following best practices
def example_function():
# Your implementation here
pass
Assumptions / Constraints / Non-goals
- Assumptions:
- Development environment is properly configured
- Required dependencies are available
- Team has basic understanding of domain
- Constraints:
- Must follow existing codebase conventions
- Time and resource limitations
- Compatibility requirements
- Non-goals:
- This skill does not cover edge cases outside scope
- Not a replacement for formal training
Compatibility & Prerequisites
- Supported Versions:
- Python 3.8+
- Node.js 16+
- Modern browsers (Chrome, Firefox, Safari, Edge)
- Required AI Tools:
- Code editor (VS Code recommended)
- Testing framework appropriate for language
- Version control (Git)
- Dependencies:
- Language-specific package manager
- Build tools
- Testing libraries
- Environment Setup:
.env.example keys: API_KEY, DATABASE_URL (no values)
Test Scenario Matrix (QA Strategy)
| Type |
Focus Area |
Required Scenarios / Mocks |
| Unit |
Core Logic |
Must cover primary logic and at least 3 edge/error cases. Target minimum 80% coverage |
| Integration |
DB / API |
All external API calls or database connections must be mocked during unit tests |
| E2E |
User Journey |
Critical user flows to test |
| Performance |
Latency / Load |
Benchmark requirements |
| Security |
Vuln / Auth |
SAST/DAST or dependency audit |
| Frontend |
UX / A11y |
Accessibility checklist (WCAG), Performance Budget (Lighthouse score) |
Technical Guardrails & Security Threat Model
1. Security & Privacy (Threat Model)
- Top Threats: Injection attacks, authentication bypass, data exposure
2. Performance & Resources
3. Architecture & Scalability
4. Observability & Reliability
Agent Directives & Error Recovery
(ข้อกำหนดสำหรับ AI Agent ในการคิดและแก้ปัญหาเมื่อเกิดข้อผิดพลาด)
- Thinking Process: Analyze root cause before fixing. Do not brute-force.
- Fallback Strategy: Stop after 3 failed test attempts. Output root cause and ask for human intervention/clarification.
- Self-Review: Check against Guardrails & Anti-patterns before finalizing.
- Output Constraints: Output ONLY the modified code block. Do not explain unless asked.
Definition of Done (DoD) Checklist
Anti-patterns
Reference Links & Examples
- Internal documentation and examples
- Official documentation and best practices
- Community resources and discussions
Versioning & Changelog
- Version: 1.0.0
- Changelog:
- 2026-02-22: Initial version with complete template structure
Converted and distributed by TomeVault — claim your Tome and manage your conversions.
1---2name: ab-testing-ml3description: A/B testing compares ML models in production. This guide covers experiment Use when this capability is needed.4---56# Ab Testing Ml78## Skill Profile9*(Select at least one profile to enable specific modules)*10- [ ] **DevOps**11- [x] **Backend**12- [ ] **Frontend**13- [ ] **AI-RAG**14- [ ] **Security Critical**1516## Overview17A/B testing compares ML models in production. This guide covers experiment design, statistical testing, multi-armed bandits, and monitoring for safely deploying and comparing ML models to determine which performs better on business metrics.1819## Why This Matters20- **Risk Mitigation**: A/B testing allows safe model deployment with gradual rollout21- **Data-Driven Decisions**: Statistical testing ensures decisions are based on significant results22- **Business Impact**: Multi-armed bandits optimize for business metrics during experimentation2324---2526## Core Concepts & Rules2728### 1. Core Principles29- Follow established patterns and conventions30- Maintain consistency across codebase31- Document decisions and trade-offs3233### 2. Implementation Guidelines34- Start with the simplest viable solution35- Iterate based on feedback and requirements36- Test thoroughly before deployment373839## Inputs / Outputs / Contracts40* **Inputs**:41 - User ID for consistent variant assignment42 - Features for model prediction43 - Actual outcomes for metric calculation44* **Entry Conditions**:45 - Two or more models ready for production46 - Sufficient traffic for statistical significance47 - Metrics tracking infrastructure in place48* **Outputs**:49 - Variant assignment (A/B)50 - Model prediction with variant info51 - Statistical test results (p-value, significance)52 - Winner recommendation53* **Artifacts Required (Deliverables)**:54 - Experiment design configuration55 - Traffic splitting logic56 - Statistical testing implementation57 - Multi-armed bandit algorithms58 - Monitoring dashboard59* **Acceptance Evidence**:60 - A/B test API endpoints functional61 - Statistical tests passing62 - Monitoring dashboard showing results63* **Success Criteria**:64 - Statistical significance (p < 0.05)65 - Minimum sample size achieved66 - Business metrics improvement validated6768## Skill Composition69* **Depends on**: ML models, Metrics tracking, Feature flags70* **Compatible with**: ML Serving, Model Experiments, Feature Engineering71* **Conflicts with**: None72* **Related Skills**: [ml-serving](39-data-science-ml/ml-serving/SKILL.md), [model-experiments](39-data-science-ml/model-experiments/SKILL.md), [feature-engineering](39-data-science-ml/feature-engineering/SKILL.md)7374---7576## Quick Start / Implementation Example77781. Review requirements and constraints792. Set up development environment803. Implement core functionality following patterns814. Write tests for critical paths825. Run tests and fix issues836. Document any deviations or decisions8485```python86# Example implementation following best practices87def example_function():88 # Your implementation here89 pass90```919293## Assumptions / Constraints / Non-goals9495* **Assumptions**:96 - Development environment is properly configured97 - Required dependencies are available98 - Team has basic understanding of domain99* **Constraints**:100 - Must follow existing codebase conventions101 - Time and resource limitations102 - Compatibility requirements103* **Non-goals**:104 - This skill does not cover edge cases outside scope105 - Not a replacement for formal training106107108## Compatibility & Prerequisites109110* **Supported Versions**:111 - Python 3.8+112 - Node.js 16+113 - Modern browsers (Chrome, Firefox, Safari, Edge)114* **Required AI Tools**:115 - Code editor (VS Code recommended)116 - Testing framework appropriate for language117 - Version control (Git)118* **Dependencies**:119 - Language-specific package manager120 - Build tools121 - Testing libraries122* **Environment Setup**:123 - `.env.example` keys: `API_KEY`, `DATABASE_URL` (no values)124125126## Test Scenario Matrix (QA Strategy)127128| Type | Focus Area | Required Scenarios / Mocks |129| :--- | :--- | :--- |130| **Unit** | Core Logic | Must cover primary logic and at least 3 edge/error cases. Target minimum 80% coverage |131| **Integration** | DB / API | All external API calls or database connections must be mocked during unit tests |132| **E2E** | User Journey | Critical user flows to test |133| **Performance** | Latency / Load | Benchmark requirements |134| **Security** | Vuln / Auth | SAST/DAST or dependency audit |135| **Frontend** | UX / A11y | Accessibility checklist (WCAG), Performance Budget (Lighthouse score) |136137138## Technical Guardrails & Security Threat Model139140### 1. Security & Privacy (Threat Model)141* **Top Threats**: Injection attacks, authentication bypass, data exposure142- [ ] **Data Handling**: Sanitize all user inputs to prevent Injection attacks. Never log raw PII143- [ ] **Secrets Management**: No hardcoded API keys. Use Env Vars/Secrets Manager144- [ ] **Authorization**: Validate user permissions before state changes145146### 2. Performance & Resources147- [ ] **Execution Efficiency**: Consider time complexity for algorithms148- [ ] **Memory Management**: Use streams/pagination for large data149- [ ] **Resource Cleanup**: Close DB connections/file handlers in finally blocks150151### 3. Architecture & Scalability152- [ ] **Design Pattern**: Follow SOLID principles, use Dependency Injection153- [ ] **Modularity**: Decouple logic from UI/Frameworks154155### 4. Observability & Reliability156- [ ] **Logging Standards**: Structured JSON, include trace IDs `request_id`157- [ ] **Metrics**: Track `error_rate`, `latency`, `queue_depth`158- [ ] **Error Handling**: Standardized error codes, no bare except159- [ ] **Observability Artifacts**:160 - **Log Fields**: timestamp, level, message, request_id161 - **Metrics**: request_count, error_count, response_time162 - **Dashboards/Alerts**: High Error Rate > 5%163164165## Agent Directives & Error Recovery166*(ข้อกำหนดสำหรับ AI Agent ในการคิดและแก้ปัญหาเมื่อเกิดข้อผิดพลาด)*167168- **Thinking Process**: Analyze root cause before fixing. Do not brute-force.169- **Fallback Strategy**: Stop after 3 failed test attempts. Output root cause and ask for human intervention/clarification.170- **Self-Review**: Check against Guardrails & Anti-patterns before finalizing.171- **Output Constraints**: Output ONLY the modified code block. Do not explain unless asked.172173174## Definition of Done (DoD) Checklist175176- [ ] Tests passed + coverage met177- [ ] Lint/Typecheck passed178- [ ] Logging/Metrics/Trace implemented179- [ ] Security checks passed180- [ ] Documentation/Changelog updated181- [ ] Accessibility/Performance requirements met (if frontend)182183184## Anti-patterns185#186187## Reference Links & Examples188189* Internal documentation and examples190* Official documentation and best practices191* Community resources and discussions192193194## Versioning & Changelog195196* **Version**: 1.0.0197* **Changelog**:198 - 2026-02-22: Initial version with complete template structure199200---201> Converted and distributed by [TomeVault](https://tomevault.io/claim/amnadtaowsoam) — claim your Tome and manage your conversions.202<!-- tomevault:4.0:skill_md:2026-04-13 -->