Model Training
Skill Profile
(Select at least one profile to enable specific modules)
Overview
Model training is the process of teaching machine learning models to make predictions or decisions based on data. This skill covers comprehensive training workflows including pipeline design, data preparation, training loops, hyperparameter tuning, experiment tracking, checkpoint management, early stopping, learning rate scheduling, distributed training, and model evaluation.
Why This Matters
- Model Quality: Proper training ensures better models
- Reproducibility: Consistent training workflows
- Efficiency: Optimized training saves time and resources
- Experimentation: Systematic hyperparameter exploration
- Production: Reliable training for deployment
Core Concepts & Rules
1. Core Principles
- Follow established patterns and conventions
- Maintain consistency across codebase
- Document decisions and trade-offs
2. Implementation Guidelines
- Start with the simplest viable solution
- Iterate based on feedback and requirements
- Test thoroughly before deployment
Inputs / Outputs / Contracts
- Inputs:
- Training data (train, validation, test)
- Model architecture
- Training configuration (hyperparameters, epochs, batch size)
- Checkpoint directory
- Entry Conditions:
- Data is properly preprocessed and split
- Model architecture is defined
- Training configuration is validated
- Compute resources (GPU) are available
- Outputs:
- Trained model
- Training metrics (loss, accuracy, etc.)
- Checkpoints
- Best model
- Experiment logs
- Artifacts Required (Deliverables):
- Trained model file (.pt, .pth)
- Training metrics logs
- Checkpoint files
- Configuration files
- Training script
- Acceptance Evidence:
- Training loss curve showing convergence
- Validation metrics meeting target thresholds
- Checkpoint files saved successfully
- Best model loaded and verified
- Success Criteria:
- Training loss converges to stable value
- Validation metrics meet or exceed targets (e.g., accuracy > 90%)
- No overfitting (train/val gap < 5%)
- Training completes within expected time
Skill Composition
- Depends on:
05-ai-ml-core/data-augmentation
05-ai-ml-core/data-preprocessing
- Compatible with:
05-ai-ml-core/model-optimization
77-mlops-data-engineering/mlflow-patterns
- Conflicts with: None
- Related Skills:
05-ai-ml-core/model-optimization
77-mlops-data-engineering/mlflow-patterns
Quick Start / Implementation Example
- Review requirements and constraints
- Set up development environment
- Implement core functionality following patterns
- Write tests for critical paths
- Run tests and fix issues
- Document any deviations or decisions
# Example implementation following best practices
def example_function():
# Your implementation here
pass
Assumptions / Constraints / Non-goals
- Assumptions:
- Development environment is properly configured
- Required dependencies are available
- Team has basic understanding of domain
- Constraints:
- Must follow existing codebase conventions
- Time and resource limitations
- Compatibility requirements
- Non-goals:
- This skill does not cover edge cases outside scope
- Not a replacement for formal training
Compatibility & Prerequisites
- Supported Versions:
- Python 3.8+
- Node.js 16+
- Modern browsers (Chrome, Firefox, Safari, Edge)
- Required AI Tools:
- Code editor (VS Code recommended)
- Testing framework appropriate for language
- Version control (Git)
- Dependencies:
- Language-specific package manager
- Build tools
- Testing libraries
- Environment Setup:
.env.example keys: API_KEY, DATABASE_URL (no values)
Test Scenario Matrix (QA Strategy)
| Type |
Focus Area |
Required Scenarios / Mocks |
| Unit |
Core Logic |
Must cover primary logic and at least 3 edge/error cases. Target minimum 80% coverage |
| Integration |
DB / API |
All external API calls or database connections must be mocked during unit tests |
| E2E |
User Journey |
Critical user flows to test |
| Performance |
Latency / Load |
Benchmark requirements |
| Security |
Vuln / Auth |
SAST/DAST or dependency audit |
| Frontend |
UX / A11y |
Accessibility checklist (WCAG), Performance Budget (Lighthouse score) |
Technical Guardrails & Security Threat Model
1. Security & Privacy (Threat Model)
- Top Threats: Injection attacks, authentication bypass, data exposure
2. Performance & Resources
3. Architecture & Scalability
4. Observability & Reliability
Agent Directives & Error Recovery
(ข้อกำหนดสำหรับ AI Agent ในการคิดและแก้ปัญหาเมื่อเกิดข้อผิดพลาด)
- Thinking Process: Analyze root cause before fixing. Do not brute-force.
- Fallback Strategy: Stop after 3 failed test attempts. Output root cause and ask for human intervention/clarification.
- Self-Review: Check against Guardrails & Anti-patterns before finalizing.
- Output Constraints: Output ONLY the modified code block. Do not explain unless asked.
Definition of Done (DoD) Checklist
Anti-patterns / Pitfalls
- ⛔ Don't: Log PII, catch-all exception, N+1 queries
- ⚠️ Watch out for: Common symptoms and quick fixes
- 💡 Instead: Use proper error handling, pagination, and logging
Reference Links & Examples
- Internal documentation and examples
- Official documentation and best practices
- Community resources and discussions
Versioning & Changelog
- Version: 1.0.0
- Changelog:
- 2026-02-22: Initial version with complete template structure
1---2name: model-training3description: Model training is the process of teaching machine learning models to make predictions or decisions based on data. This skill covers comprehensive training workflows including pipeline design, data pre4---5
6# Model Training
7
8## Skill Profile
9*(Select at least one profile to enable specific modules)*
10- [ ] **DevOps**
11- [x] **Backend**
12- [ ] **Frontend**
13- [ ] **AI-RAG**
14- [ ] **Security Critical**
15
16## Overview
17Model training is the process of teaching machine learning models to make predictions or decisions based on data. This skill covers comprehensive training workflows including pipeline design, data preparation, training loops, hyperparameter tuning, experiment tracking, checkpoint management, early stopping, learning rate scheduling, distributed training, and model evaluation.
18
19## Why This Matters
20- **Model Quality**: Proper training ensures better models
21- **Reproducibility**: Consistent training workflows
22- **Efficiency**: Optimized training saves time and resources
23- **Experimentation**: Systematic hyperparameter exploration
24- **Production**: Reliable training for deployment
25
26---
27
28## Core Concepts & Rules
29
30### 1. Core Principles
31- Follow established patterns and conventions
32- Maintain consistency across codebase
33- Document decisions and trade-offs
34
35### 2. Implementation Guidelines
36- Start with the simplest viable solution
37- Iterate based on feedback and requirements
38- Test thoroughly before deployment
39
40
41## Inputs / Outputs / Contracts
42* **Inputs**:
43 - Training data (train, validation, test)
44 - Model architecture
45 - Training configuration (hyperparameters, epochs, batch size)
46 - Checkpoint directory
47* **Entry Conditions**:
48 - Data is properly preprocessed and split
49 - Model architecture is defined
50 - Training configuration is validated
51 - Compute resources (GPU) are available
52* **Outputs**:
53 - Trained model
54 - Training metrics (loss, accuracy, etc.)
55 - Checkpoints
56 - Best model
57 - Experiment logs
58* **Artifacts Required (Deliverables)**:
59 - Trained model file (.pt, .pth)
60 - Training metrics logs
61 - Checkpoint files
62 - Configuration files
63 - Training script
64* **Acceptance Evidence**:
65 - Training loss curve showing convergence
66 - Validation metrics meeting target thresholds
67 - Checkpoint files saved successfully
68 - Best model loaded and verified
69* **Success Criteria**:
70 - Training loss converges to stable value
71 - Validation metrics meet or exceed targets (e.g., accuracy > 90%)
72 - No overfitting (train/val gap < 5%)
73 - Training completes within expected time
74
75## Skill Composition
76* **Depends on**:
77 - [`05-ai-ml-core/data-augmentation`](05-ai-ml-core/data-augmentation/SKILL.md)
78 - [`05-ai-ml-core/data-preprocessing`](05-ai-ml-core/data-preprocessing/SKILL.md)
79* **Compatible with**:
80 - [`05-ai-ml-core/model-optimization`](05-ai-ml-core/model-optimization/SKILL.md)
81 - [`77-mlops-data-engineering/mlflow-patterns`](77-mlops-data-engineering/mlflow-patterns/SKILL.md)
82* **Conflicts with**: None
83* **Related Skills**:
84 - [`05-ai-ml-core/model-optimization`](05-ai-ml-core/model-optimization/SKILL.md)
85 - [`77-mlops-data-engineering/mlflow-patterns`](77-mlops-data-engineering/mlflow-patterns/SKILL.md)
86
87---
88
89## Quick Start / Implementation Example
90
911. Review requirements and constraints
922. Set up development environment
933. Implement core functionality following patterns
944. Write tests for critical paths
955. Run tests and fix issues
966. Document any deviations or decisions
97
98```python
99# Example implementation following best practices
100def example_function():
101 # Your implementation here
102 pass
103```
104
105
106## Assumptions / Constraints / Non-goals
107
108* **Assumptions**:
109 - Development environment is properly configured
110 - Required dependencies are available
111 - Team has basic understanding of domain
112* **Constraints**:
113 - Must follow existing codebase conventions
114 - Time and resource limitations
115 - Compatibility requirements
116* **Non-goals**:
117 - This skill does not cover edge cases outside scope
118 - Not a replacement for formal training
119
120
121## Compatibility & Prerequisites
122
123* **Supported Versions**:
124 - Python 3.8+
125 - Node.js 16+
126 - Modern browsers (Chrome, Firefox, Safari, Edge)
127* **Required AI Tools**:
128 - Code editor (VS Code recommended)
129 - Testing framework appropriate for language
130 - Version control (Git)
131* **Dependencies**:
132 - Language-specific package manager
133 - Build tools
134 - Testing libraries
135* **Environment Setup**:
136 - `.env.example` keys: `API_KEY`, `DATABASE_URL` (no values)
137
138
139## Test Scenario Matrix (QA Strategy)
140
141| Type | Focus Area | Required Scenarios / Mocks |
142| :--- | :--- | :--- |
143| **Unit** | Core Logic | Must cover primary logic and at least 3 edge/error cases. Target minimum 80% coverage |
144| **Integration** | DB / API | All external API calls or database connections must be mocked during unit tests |
145| **E2E** | User Journey | Critical user flows to test |
146| **Performance** | Latency / Load | Benchmark requirements |
147| **Security** | Vuln / Auth | SAST/DAST or dependency audit |
148| **Frontend** | UX / A11y | Accessibility checklist (WCAG), Performance Budget (Lighthouse score) |
149
150
151## Technical Guardrails & Security Threat Model
152
153### 1. Security & Privacy (Threat Model)
154* **Top Threats**: Injection attacks, authentication bypass, data exposure
155- [ ] **Data Handling**: Sanitize all user inputs to prevent Injection attacks. Never log raw PII
156- [ ] **Secrets Management**: No hardcoded API keys. Use Env Vars/Secrets Manager
157- [ ] **Authorization**: Validate user permissions before state changes
158
159### 2. Performance & Resources
160- [ ] **Execution Efficiency**: Consider time complexity for algorithms
161- [ ] **Memory Management**: Use streams/pagination for large data
162- [ ] **Resource Cleanup**: Close DB connections/file handlers in finally blocks
163
164### 3. Architecture & Scalability
165- [ ] **Design Pattern**: Follow SOLID principles, use Dependency Injection
166- [ ] **Modularity**: Decouple logic from UI/Frameworks
167
168### 4. Observability & Reliability
169- [ ] **Logging Standards**: Structured JSON, include trace IDs `request_id`
170- [ ] **Metrics**: Track `error_rate`, `latency`, `queue_depth`
171- [ ] **Error Handling**: Standardized error codes, no bare except
172- [ ] **Observability Artifacts**:
173 - **Log Fields**: timestamp, level, message, request_id
174 - **Metrics**: request_count, error_count, response_time
175 - **Dashboards/Alerts**: High Error Rate > 5%
176
177
178## Agent Directives & Error Recovery
179*(ข้อกำหนดสำหรับ AI Agent ในการคิดและแก้ปัญหาเมื่อเกิดข้อผิดพลาด)*
180
181- **Thinking Process**: Analyze root cause before fixing. Do not brute-force.
182- **Fallback Strategy**: Stop after 3 failed test attempts. Output root cause and ask for human intervention/clarification.
183- **Self-Review**: Check against Guardrails & Anti-patterns before finalizing.
184- **Output Constraints**: Output ONLY the modified code block. Do not explain unless asked.
185
186
187## Definition of Done (DoD) Checklist
188
189- [ ] Tests passed + coverage met
190- [ ] Lint/Typecheck passed
191- [ ] Logging/Metrics/Trace implemented
192- [ ] Security checks passed
193- [ ] Documentation/Changelog updated
194- [ ] Accessibility/Performance requirements met (if frontend)
195
196
197## Anti-patterns / Pitfalls
198
199* ⛔ **Don't**: Log PII, catch-all exception, N+1 queries
200* ⚠️ **Watch out for**: Common symptoms and quick fixes
201* 💡 **Instead**: Use proper error handling, pagination, and logging
202
203
204## Reference Links & Examples
205
206* Internal documentation and examples
207* Official documentation and best practices
208* Community resources and discussions
209
210
211## Versioning & Changelog
212
213* **Version**: 1.0.0
214* **Changelog**:
215 - 2026-02-22: Initial version with complete template structure
216