Model Experiments
Skill Profile
(Select at least one profile to enable specific modules)
Overview
Experiment tracking manages ML experiments, metrics, parameters, and artifacts. This guide covers MLflow, Weights & Biases, and best practices for tracking experiments, comparing models, and ensuring reproducibility in ML development.
Why This Matters
- Reproducibility: Experiment tracking ensures experiments can be reproduced
- Comparison: Systematic tracking enables fair model comparison
- Collaboration: Shared experiment tracking improves team collaboration
- Model Versioning: Track model evolution and deployment history
Core Concepts & Rules
1. Core Principles
- Follow established patterns and conventions
- Maintain consistency across codebase
- Document decisions and trade-offs
2. Implementation Guidelines
- Start with the simplest viable solution
- Iterate based on feedback and requirements
- Test thoroughly before deployment
Inputs / Outputs / Contracts
- Inputs:
- Training data and validation data
- Model configuration and hyperparameters
- Experiment metadata (name, description)
- Entry Conditions:
- Experiment tracking tool installed (MLflow, W&B)
- Model training code instrumented for logging
- Storage backend configured
- Outputs:
- Experiment records with parameters and metrics
- Model artifacts saved and versioned
- Comparison dashboards
- Reproduction scripts
- Artifacts Required (Deliverables):
- Experiment tracking configuration
- Logging instrumentation in training code
- Comparison dashboards
- Model registry setup
- Acceptance Evidence:
- Experiments tracked with all metadata
- Models reproducible from logged information
- Comparison shows clear winners
- Success Criteria:
- All experiments logged with complete metadata
- Models can be reproduced
- Clear performance comparison available
Skill Composition
- Depends on: Model training code, Experiment tracking tool
- Compatible with: Feature Engineering, Model Training, ML Serving
- Conflicts with: None
- Related Skills: feature-engineering, model-training, ml-serving
Quick Start / Implementation Example
- Review requirements and constraints
- Set up development environment
- Implement core functionality following patterns
- Write tests for critical paths
- Run tests and fix issues
- Document any deviations or decisions
# Example implementation following best practices
def example_function():
# Your implementation here
pass
Assumptions / Constraints / Non-goals
- Assumptions:
- Development environment is properly configured
- Required dependencies are available
- Team has basic understanding of domain
- Constraints:
- Must follow existing codebase conventions
- Time and resource limitations
- Compatibility requirements
- Non-goals:
- This skill does not cover edge cases outside scope
- Not a replacement for formal training
Compatibility & Prerequisites
- Supported Versions:
- Python 3.8+
- Node.js 16+
- Modern browsers (Chrome, Firefox, Safari, Edge)
- Required AI Tools:
- Code editor (VS Code recommended)
- Testing framework appropriate for language
- Version control (Git)
- Dependencies:
- Language-specific package manager
- Build tools
- Testing libraries
- Environment Setup:
.env.example keys: API_KEY, DATABASE_URL (no values)
Test Scenario Matrix (QA Strategy)
| Type |
Focus Area |
Required Scenarios / Mocks |
| Unit |
Core Logic |
Must cover primary logic and at least 3 edge/error cases. Target minimum 80% coverage |
| Integration |
DB / API |
All external API calls or database connections must be mocked during unit tests |
| E2E |
User Journey |
Critical user flows to test |
| Performance |
Latency / Load |
Benchmark requirements |
| Security |
Vuln / Auth |
SAST/DAST or dependency audit |
| Frontend |
UX / A11y |
Accessibility checklist (WCAG), Performance Budget (Lighthouse score) |
Technical Guardrails & Security Threat Model
1. Security & Privacy (Threat Model)
- Top Threats: Injection attacks, authentication bypass, data exposure
2. Performance & Resources
3. Architecture & Scalability
4. Observability & Reliability
Agent Directives & Error Recovery
(ข้อกำหนดสำหรับ AI Agent ในการคิดและแก้ปัญหาเมื่อเกิดข้อผิดพลาด)
- Thinking Process: Analyze root cause before fixing. Do not brute-force.
- Fallback Strategy: Stop after 3 failed test attempts. Output root cause and ask for human intervention/clarification.
- Self-Review: Check against Guardrails & Anti-patterns before finalizing.
- Output Constraints: Output ONLY the modified code block. Do not explain unless asked.
Definition of Done (DoD) Checklist
Anti-patterns
Reference Links & Examples
- Internal documentation and examples
- Official documentation and best practices
- Community resources and discussions
Versioning & Changelog
- Version: 1.0.0
- Changelog:
- 2026-02-22: Initial version with complete template structure
Converted and distributed by TomeVault — claim your Tome and manage your conversions.
1---2name: model-experiments3description: Experiment tracking manages ML experiments, metrics, parameters, and Use when this capability is needed.4---56# Model Experiments78## Skill Profile9*(Select at least one profile to enable specific modules)*10- [ ] **DevOps**11- [x] **Backend**12- [ ] **Frontend**13- [ ] **AI-RAG**14- [ ] **Security Critical**1516## Overview17Experiment tracking manages ML experiments, metrics, parameters, and artifacts. This guide covers MLflow, Weights & Biases, and best practices for tracking experiments, comparing models, and ensuring reproducibility in ML development.1819## Why This Matters20- **Reproducibility**: Experiment tracking ensures experiments can be reproduced21- **Comparison**: Systematic tracking enables fair model comparison22- **Collaboration**: Shared experiment tracking improves team collaboration23- **Model Versioning**: Track model evolution and deployment history2425---2627## Core Concepts & Rules2829### 1. Core Principles30- Follow established patterns and conventions31- Maintain consistency across codebase32- Document decisions and trade-offs3334### 2. Implementation Guidelines35- Start with the simplest viable solution36- Iterate based on feedback and requirements37- Test thoroughly before deployment383940## Inputs / Outputs / Contracts41* **Inputs**:42 - Training data and validation data43 - Model configuration and hyperparameters44 - Experiment metadata (name, description)45* **Entry Conditions**:46 - Experiment tracking tool installed (MLflow, W&B)47 - Model training code instrumented for logging48 - Storage backend configured49* **Outputs**:50 - Experiment records with parameters and metrics51 - Model artifacts saved and versioned52 - Comparison dashboards53 - Reproduction scripts54* **Artifacts Required (Deliverables)**:55 - Experiment tracking configuration56 - Logging instrumentation in training code57 - Comparison dashboards58 - Model registry setup59* **Acceptance Evidence**:60 - Experiments tracked with all metadata61 - Models reproducible from logged information62 - Comparison shows clear winners63* **Success Criteria**:64 - All experiments logged with complete metadata65 - Models can be reproduced66 - Clear performance comparison available6768## Skill Composition69* **Depends on**: Model training code, Experiment tracking tool70* **Compatible with**: Feature Engineering, Model Training, ML Serving71* **Conflicts with**: None72* **Related Skills**: [feature-engineering](39-data-science-ml/feature-engineering/SKILL.md), [model-training](05-ai-ml-core/model-training/SKILL.md), [ml-serving](39-data-science-ml/ml-serving/SKILL.md)7374---7576## Quick Start / Implementation Example77781. Review requirements and constraints792. Set up development environment803. Implement core functionality following patterns814. Write tests for critical paths825. Run tests and fix issues836. Document any deviations or decisions8485```python86# Example implementation following best practices87def example_function():88 # Your implementation here89 pass90```919293## Assumptions / Constraints / Non-goals9495* **Assumptions**:96 - Development environment is properly configured97 - Required dependencies are available98 - Team has basic understanding of domain99* **Constraints**:100 - Must follow existing codebase conventions101 - Time and resource limitations102 - Compatibility requirements103* **Non-goals**:104 - This skill does not cover edge cases outside scope105 - Not a replacement for formal training106107108## Compatibility & Prerequisites109110* **Supported Versions**:111 - Python 3.8+112 - Node.js 16+113 - Modern browsers (Chrome, Firefox, Safari, Edge)114* **Required AI Tools**:115 - Code editor (VS Code recommended)116 - Testing framework appropriate for language117 - Version control (Git)118* **Dependencies**:119 - Language-specific package manager120 - Build tools121 - Testing libraries122* **Environment Setup**:123 - `.env.example` keys: `API_KEY`, `DATABASE_URL` (no values)124125126## Test Scenario Matrix (QA Strategy)127128| Type | Focus Area | Required Scenarios / Mocks |129| :--- | :--- | :--- |130| **Unit** | Core Logic | Must cover primary logic and at least 3 edge/error cases. Target minimum 80% coverage |131| **Integration** | DB / API | All external API calls or database connections must be mocked during unit tests |132| **E2E** | User Journey | Critical user flows to test |133| **Performance** | Latency / Load | Benchmark requirements |134| **Security** | Vuln / Auth | SAST/DAST or dependency audit |135| **Frontend** | UX / A11y | Accessibility checklist (WCAG), Performance Budget (Lighthouse score) |136137138## Technical Guardrails & Security Threat Model139140### 1. Security & Privacy (Threat Model)141* **Top Threats**: Injection attacks, authentication bypass, data exposure142- [ ] **Data Handling**: Sanitize all user inputs to prevent Injection attacks. Never log raw PII143- [ ] **Secrets Management**: No hardcoded API keys. Use Env Vars/Secrets Manager144- [ ] **Authorization**: Validate user permissions before state changes145146### 2. Performance & Resources147- [ ] **Execution Efficiency**: Consider time complexity for algorithms148- [ ] **Memory Management**: Use streams/pagination for large data149- [ ] **Resource Cleanup**: Close DB connections/file handlers in finally blocks150151### 3. Architecture & Scalability152- [ ] **Design Pattern**: Follow SOLID principles, use Dependency Injection153- [ ] **Modularity**: Decouple logic from UI/Frameworks154155### 4. Observability & Reliability156- [ ] **Logging Standards**: Structured JSON, include trace IDs `request_id`157- [ ] **Metrics**: Track `error_rate`, `latency`, `queue_depth`158- [ ] **Error Handling**: Standardized error codes, no bare except159- [ ] **Observability Artifacts**:160 - **Log Fields**: timestamp, level, message, request_id161 - **Metrics**: request_count, error_count, response_time162 - **Dashboards/Alerts**: High Error Rate > 5%163164165## Agent Directives & Error Recovery166*(ข้อกำหนดสำหรับ AI Agent ในการคิดและแก้ปัญหาเมื่อเกิดข้อผิดพลาด)*167168- **Thinking Process**: Analyze root cause before fixing. Do not brute-force.169- **Fallback Strategy**: Stop after 3 failed test attempts. Output root cause and ask for human intervention/clarification.170- **Self-Review**: Check against Guardrails & Anti-patterns before finalizing.171- **Output Constraints**: Output ONLY the modified code block. Do not explain unless asked.172173174## Definition of Done (DoD) Checklist175176- [ ] Tests passed + coverage met177- [ ] Lint/Typecheck passed178- [ ] Logging/Metrics/Trace implemented179- [ ] Security checks passed180- [ ] Documentation/Changelog updated181- [ ] Accessibility/Performance requirements met (if frontend)182183184## Anti-patterns185#186187## Reference Links & Examples188189* Internal documentation and examples190* Official documentation and best practices191* Community resources and discussions192193194## Versioning & Changelog195196* **Version**: 1.0.0197* **Changelog**:198 - 2026-02-22: Initial version with complete template structure199200---201> Converted and distributed by [TomeVault](https://tomevault.io/claim/amnadtaowsoam) — claim your Tome and manage your conversions.202<!-- tomevault:4.0:skill_md:2026-04-13 -->