1---2name: fine-tuning3description: Fine-tune LLMs with data preparation, provider selection, cost estimation, evaluation, and compliance checks.4---5
6## When to Use
7
8User wants to fine-tune a language model, evaluate if fine-tuning is worth it, or debug training issues.
9
10## Quick Reference
11
12| Topic | File |
13|-------|------|
14| Provider comparison & pricing | `providers.md` |
15| Data preparation & validation | `data-prep.md` |
16| Training configuration | `training.md` |
17| Evaluation & debugging | `evaluation.md` |
18| Cost estimation & ROI | `costs.md` |
19| Compliance & security | `compliance.md` |
20
21## Core Capabilities
22
231. **Decide fit** — Analyze if fine-tuning beats prompting for the use case
242. **Prepare data** — Convert raw data to JSONL, deduplicate, validate format
253. **Select provider** — Compare OpenAI, Anthropic (Bedrock), Google, open source based on constraints
264. **Estimate costs** — Calculate training cost, inference savings, break-even point
275. **Configure training** — Set hyperparameters (learning rate, epochs, LoRA rank)
286. **Run evaluation** — Compare fine-tuned vs base model on task-specific metrics
297. **Debug failures** — Diagnose loss curves, overfitting, catastrophic forgetting
308. **Handle compliance** — Scan for PII, configure on-premise training, generate audit logs
31
32## Decision Checklist
33
34Before recommending fine-tuning, ask:
35- [ ] What's the failure mode with prompting? (format, style, knowledge, cost)
36- [ ] How many training examples available? (minimum 50-100)
37- [ ] Expected inference volume? (affects ROI calculation)
38- [ ] Privacy constraints? (determines provider options)
39- [ ] Budget for training + ongoing inference?
40
41## Fine-Tune vs Prompt Decision
42
43| Signal | Recommendation |
44|--------|----------------|
45| Format/style inconsistency | Fine-tune ✓ |
46| Missing domain knowledge | RAG first, then fine-tune if needed |
47| High inference volume (>100K/mo) | Fine-tune for cost savings |
48| Requirements change frequently | Stick with prompting |
49| <50 quality examples | Prompting + few-shot |
50
51## Critical Rules
52
53- **Data quality > quantity** — 100 great examples beat 1000 noisy ones
54- **LoRA first** — Never jump to full fine-tuning; LoRA is 10-100x cheaper
55- **Hold out eval set** — Always 80/10/10 split; never peek at test data
56- **Same precision** — Train and serve at identical precision (4-bit, 16-bit)
57- **Baseline first** — Run eval on base model before training to measure actual improvement
58- **Expect iteration** — First attempt rarely optimal; plan for 2-3 cycles
59
60## Common Pitfalls
61
62| Mistake | Fix |
63|---------|-----|
64| Training on inconsistent data | Manual review of 100+ samples before training |
65| Learning rate too high | Start with 2e-4 for SFT, 5e-6 for RLHF |
66| Expecting new knowledge | Fine-tuning adjusts behavior, not knowledge — use RAG |
67| No baseline comparison | Always test base model on same eval set |
68| Ignoring forgetting | Mix 20% general data to preserve capabilities |