Model Training
Skill Profile
(Select at least one profile to enable specific modules)
Overview
Model training is the process of teaching machine learning models to make predictions or decisions based on data. This skill covers comprehensive training workflows including pipeline design, data preparation, training loops, hyperparameter tuning, experiment tracking, checkpoint management, early stopping, learning rate scheduling, distributed training, and model evaluation.
Why This Matters
- Model Quality: Proper training ensures better models
- Reproducibility: Consistent training workflows
- Efficiency: Optimized training saves time and resources
- Experimentation: Systematic hyperparameter exploration
- Production: Reliable training for deployment
Core Concepts & Rules
1. Core Principles
- Follow established patterns and conventions
- Maintain consistency across codebase
- Document decisions and trade-offs
2. Implementation Guidelines
- Start with the simplest viable solution
- Iterate based on feedback and requirements
- Test thoroughly before deployment
Inputs / Outputs / Contracts
- Inputs:
- Training data (train, validation, test)
- Model architecture
- Training configuration (hyperparameters, epochs, batch size)
- Checkpoint directory
- Entry Conditions:
- Data is properly preprocessed and split
- Model architecture is defined
- Training configuration is validated
- Compute resources (GPU) are available
- Outputs:
- Trained model
- Training metrics (loss, accuracy, etc.)
- Checkpoints
- Best model
- Experiment logs
- Artifacts Required (Deliverables):
- Trained model file (.pt, .pth)
- Training metrics logs
- Checkpoint files
- Configuration files
- Training script
- Acceptance Evidence:
- Training loss curve showing convergence
- Validation metrics meeting target thresholds
- Checkpoint files saved successfully
- Best model loaded and verified
- Success Criteria:
- Training loss converges to stable value
- Validation metrics meet or exceed targets (e.g., accuracy > 90%)
- No overfitting (train/val gap < 5%)
- Training completes within expected time
Skill Composition
- Depends on:
05-ai-ml-core/data-augmentation
05-ai-ml-core/data-preprocessing
- Compatible with:
05-ai-ml-core/model-optimization
77-mlops-data-engineering/mlflow-patterns
- Conflicts with: None
- Related Skills:
05-ai-ml-core/model-optimization
77-mlops-data-engineering/mlflow-patterns
Quick Start / Implementation Example
- Review requirements and constraints
- Set up development environment
- Implement core functionality following patterns
- Write tests for critical paths
- Run tests and fix issues
- Document any deviations or decisions
# Example implementation following best practices
def example_function():
# Your implementation here
pass
Assumptions / Constraints / Non-goals
- Assumptions:
- Development environment is properly configured
- Required dependencies are available
- Team has basic understanding of domain
- Constraints:
- Must follow existing codebase conventions
- Time and resource limitations
- Compatibility requirements
- Non-goals:
- This skill does not cover edge cases outside scope
- Not a replacement for formal training
Compatibility & Prerequisites
- Supported Versions:
- Python 3.8+
- Node.js 16+
- Modern browsers (Chrome, Firefox, Safari, Edge)
- Required AI Tools:
- Code editor (VS Code recommended)
- Testing framework appropriate for language
- Version control (Git)
- Dependencies:
- Language-specific package manager
- Build tools
- Testing libraries
- Environment Setup:
.env.example keys: API_KEY, DATABASE_URL (no values)
Test Scenario Matrix (QA Strategy)
| Type |
Focus Area |
Required Scenarios / Mocks |
| Unit |
Core Logic |
Must cover primary logic and at least 3 edge/error cases. Target minimum 80% coverage |
| Integration |
DB / API |
All external API calls or database connections must be mocked during unit tests |
| E2E |
User Journey |
Critical user flows to test |
| Performance |
Latency / Load |
Benchmark requirements |
| Security |
Vuln / Auth |
SAST/DAST or dependency audit |
| Frontend |
UX / A11y |
Accessibility checklist (WCAG), Performance Budget (Lighthouse score) |
Technical Guardrails & Security Threat Model
1. Security & Privacy (Threat Model)
- Top Threats: Injection attacks, authentication bypass, data exposure
2. Performance & Resources
3. Architecture & Scalability
4. Observability & Reliability
Agent Directives & Error Recovery
(ข้อกำหนดสำหรับ AI Agent ในการคิดและแก้ปัญหาเมื่อเกิดข้อผิดพลาด)
- Thinking Process: Analyze root cause before fixing. Do not brute-force.
- Fallback Strategy: Stop after 3 failed test attempts. Output root cause and ask for human intervention/clarification.
- Self-Review: Check against Guardrails & Anti-patterns before finalizing.
- Output Constraints: Output ONLY the modified code block. Do not explain unless asked.
Definition of Done (DoD) Checklist
Anti-patterns / Pitfalls
- ⛔ Don't: Log PII, catch-all exception, N+1 queries
- ⚠️ Watch out for: Common symptoms and quick fixes
- 💡 Instead: Use proper error handling, pagination, and logging
Reference Links & Examples
- Internal documentation and examples
- Official documentation and best practices
- Community resources and discussions
Versioning & Changelog
- Version: 1.0.0
- Changelog:
- 2026-02-22: Initial version with complete template structure
Converted and distributed by TomeVault — claim your Tome and manage your conversions.
1---2name: model-training3description: Model training is the process of teaching machine learning models to Use when this capability is needed.4---56# Model Training78## Skill Profile9*(Select at least one profile to enable specific modules)*10- [ ] **DevOps**11- [x] **Backend**12- [ ] **Frontend**13- [ ] **AI-RAG**14- [ ] **Security Critical**1516## Overview17Model training is the process of teaching machine learning models to make predictions or decisions based on data. This skill covers comprehensive training workflows including pipeline design, data preparation, training loops, hyperparameter tuning, experiment tracking, checkpoint management, early stopping, learning rate scheduling, distributed training, and model evaluation.1819## Why This Matters20- **Model Quality**: Proper training ensures better models21- **Reproducibility**: Consistent training workflows22- **Efficiency**: Optimized training saves time and resources23- **Experimentation**: Systematic hyperparameter exploration24- **Production**: Reliable training for deployment2526---2728## Core Concepts & Rules2930### 1. Core Principles31- Follow established patterns and conventions32- Maintain consistency across codebase33- Document decisions and trade-offs3435### 2. Implementation Guidelines36- Start with the simplest viable solution37- Iterate based on feedback and requirements38- Test thoroughly before deployment394041## Inputs / Outputs / Contracts42* **Inputs**:43 - Training data (train, validation, test)44 - Model architecture45 - Training configuration (hyperparameters, epochs, batch size)46 - Checkpoint directory47* **Entry Conditions**:48 - Data is properly preprocessed and split49 - Model architecture is defined50 - Training configuration is validated51 - Compute resources (GPU) are available52* **Outputs**:53 - Trained model54 - Training metrics (loss, accuracy, etc.)55 - Checkpoints56 - Best model57 - Experiment logs58* **Artifacts Required (Deliverables)**:59 - Trained model file (.pt, .pth)60 - Training metrics logs61 - Checkpoint files62 - Configuration files63 - Training script64* **Acceptance Evidence**:65 - Training loss curve showing convergence66 - Validation metrics meeting target thresholds67 - Checkpoint files saved successfully68 - Best model loaded and verified69* **Success Criteria**:70 - Training loss converges to stable value71 - Validation metrics meet or exceed targets (e.g., accuracy > 90%)72 - No overfitting (train/val gap < 5%)73 - Training completes within expected time7475## Skill Composition76* **Depends on**: 77 - [`05-ai-ml-core/data-augmentation`](05-ai-ml-core/data-augmentation/SKILL.md)78 - [`05-ai-ml-core/data-preprocessing`](05-ai-ml-core/data-preprocessing/SKILL.md)79* **Compatible with**: 80 - [`05-ai-ml-core/model-optimization`](05-ai-ml-core/model-optimization/SKILL.md)81 - [`77-mlops-data-engineering/mlflow-patterns`](77-mlops-data-engineering/mlflow-patterns/SKILL.md)82* **Conflicts with**: None83* **Related Skills**: 84 - [`05-ai-ml-core/model-optimization`](05-ai-ml-core/model-optimization/SKILL.md)85 - [`77-mlops-data-engineering/mlflow-patterns`](77-mlops-data-engineering/mlflow-patterns/SKILL.md)8687---8889## Quick Start / Implementation Example90911. Review requirements and constraints922. Set up development environment933. Implement core functionality following patterns944. Write tests for critical paths955. Run tests and fix issues966. Document any deviations or decisions9798```python99# Example implementation following best practices100def example_function():101 # Your implementation here102 pass103```104105106## Assumptions / Constraints / Non-goals107108* **Assumptions**:109 - Development environment is properly configured110 - Required dependencies are available111 - Team has basic understanding of domain112* **Constraints**:113 - Must follow existing codebase conventions114 - Time and resource limitations115 - Compatibility requirements116* **Non-goals**:117 - This skill does not cover edge cases outside scope118 - Not a replacement for formal training119120121## Compatibility & Prerequisites122123* **Supported Versions**:124 - Python 3.8+125 - Node.js 16+126 - Modern browsers (Chrome, Firefox, Safari, Edge)127* **Required AI Tools**:128 - Code editor (VS Code recommended)129 - Testing framework appropriate for language130 - Version control (Git)131* **Dependencies**:132 - Language-specific package manager133 - Build tools134 - Testing libraries135* **Environment Setup**:136 - `.env.example` keys: `API_KEY`, `DATABASE_URL` (no values)137138139## Test Scenario Matrix (QA Strategy)140141| Type | Focus Area | Required Scenarios / Mocks |142| :--- | :--- | :--- |143| **Unit** | Core Logic | Must cover primary logic and at least 3 edge/error cases. Target minimum 80% coverage |144| **Integration** | DB / API | All external API calls or database connections must be mocked during unit tests |145| **E2E** | User Journey | Critical user flows to test |146| **Performance** | Latency / Load | Benchmark requirements |147| **Security** | Vuln / Auth | SAST/DAST or dependency audit |148| **Frontend** | UX / A11y | Accessibility checklist (WCAG), Performance Budget (Lighthouse score) |149150151## Technical Guardrails & Security Threat Model152153### 1. Security & Privacy (Threat Model)154* **Top Threats**: Injection attacks, authentication bypass, data exposure155- [ ] **Data Handling**: Sanitize all user inputs to prevent Injection attacks. Never log raw PII156- [ ] **Secrets Management**: No hardcoded API keys. Use Env Vars/Secrets Manager157- [ ] **Authorization**: Validate user permissions before state changes158159### 2. Performance & Resources160- [ ] **Execution Efficiency**: Consider time complexity for algorithms161- [ ] **Memory Management**: Use streams/pagination for large data162- [ ] **Resource Cleanup**: Close DB connections/file handlers in finally blocks163164### 3. Architecture & Scalability165- [ ] **Design Pattern**: Follow SOLID principles, use Dependency Injection166- [ ] **Modularity**: Decouple logic from UI/Frameworks167168### 4. Observability & Reliability169- [ ] **Logging Standards**: Structured JSON, include trace IDs `request_id`170- [ ] **Metrics**: Track `error_rate`, `latency`, `queue_depth`171- [ ] **Error Handling**: Standardized error codes, no bare except172- [ ] **Observability Artifacts**:173 - **Log Fields**: timestamp, level, message, request_id174 - **Metrics**: request_count, error_count, response_time175 - **Dashboards/Alerts**: High Error Rate > 5%176177178## Agent Directives & Error Recovery179*(ข้อกำหนดสำหรับ AI Agent ในการคิดและแก้ปัญหาเมื่อเกิดข้อผิดพลาด)*180181- **Thinking Process**: Analyze root cause before fixing. Do not brute-force.182- **Fallback Strategy**: Stop after 3 failed test attempts. Output root cause and ask for human intervention/clarification.183- **Self-Review**: Check against Guardrails & Anti-patterns before finalizing.184- **Output Constraints**: Output ONLY the modified code block. Do not explain unless asked.185186187## Definition of Done (DoD) Checklist188189- [ ] Tests passed + coverage met190- [ ] Lint/Typecheck passed191- [ ] Logging/Metrics/Trace implemented192- [ ] Security checks passed193- [ ] Documentation/Changelog updated194- [ ] Accessibility/Performance requirements met (if frontend)195196197## Anti-patterns / Pitfalls198199* ⛔ **Don't**: Log PII, catch-all exception, N+1 queries200* ⚠️ **Watch out for**: Common symptoms and quick fixes201* 💡 **Instead**: Use proper error handling, pagination, and logging202203204## Reference Links & Examples205206* Internal documentation and examples207* Official documentation and best practices208* Community resources and discussions209210211## Versioning & Changelog212213* **Version**: 1.0.0214* **Changelog**:215 - 2026-02-22: Initial version with complete template structure216217---218> Converted and distributed by [TomeVault](https://tomevault.io/claim/amnadtaowsoam) — claim your Tome and manage your conversions.219<!-- tomevault:4.0:skill_md:2026-04-13 -->