1---2name: fine-tuning3description: Fine-tune LLMs with data preparation, provider selection, cost estimation, evaluation, and compliance checks.4---56## When to Use78User wants to fine-tune a language model, evaluate if fine-tuning is worth it, or debug training issues.910## Quick Reference1112| Topic | File |13|-------|------|14| Provider comparison & pricing | `providers.md` |15| Data preparation & validation | `data-prep.md` |16| Training configuration | `training.md` |17| Evaluation & debugging | `evaluation.md` |18| Cost estimation & ROI | `costs.md` |19| Compliance & security | `compliance.md` |2021## Core Capabilities22231. **Decide fit** — Analyze if fine-tuning beats prompting for the use case242. **Prepare data** — Convert raw data to JSONL, deduplicate, validate format253. **Select provider** — Compare OpenAI, Anthropic (Bedrock), Google, open source based on constraints264. **Estimate costs** — Calculate training cost, inference savings, break-even point275. **Configure training** — Set hyperparameters (learning rate, epochs, LoRA rank)286. **Run evaluation** — Compare fine-tuned vs base model on task-specific metrics297. **Debug failures** — Diagnose loss curves, overfitting, catastrophic forgetting308. **Handle compliance** — Scan for PII, configure on-premise training, generate audit logs3132## Decision Checklist3334Before recommending fine-tuning, ask:35- [ ] What's the failure mode with prompting? (format, style, knowledge, cost)36- [ ] How many training examples available? (minimum 50-100)37- [ ] Expected inference volume? (affects ROI calculation)38- [ ] Privacy constraints? (determines provider options)39- [ ] Budget for training + ongoing inference?4041## Fine-Tune vs Prompt Decision4243| Signal | Recommendation |44|--------|----------------|45| Format/style inconsistency | Fine-tune ✓ |46| Missing domain knowledge | RAG first, then fine-tune if needed |47| High inference volume (>100K/mo) | Fine-tune for cost savings |48| Requirements change frequently | Stick with prompting |49| <50 quality examples | Prompting + few-shot |5051## Critical Rules5253- **Data quality > quantity** — 100 great examples beat 1000 noisy ones54- **LoRA first** — Never jump to full fine-tuning; LoRA is 10-100x cheaper55- **Hold out eval set** — Always 80/10/10 split; never peek at test data56- **Same precision** — Train and serve at identical precision (4-bit, 16-bit)57- **Baseline first** — Run eval on base model before training to measure actual improvement58- **Expect iteration** — First attempt rarely optimal; plan for 2-3 cycles5960## Common Pitfalls6162| Mistake | Fix |63|---------|-----|64| Training on inconsistent data | Manual review of 100+ samples before training |65| Learning rate too high | Start with 2e-4 for SFT, 5e-6 for RLHF |66| Expecting new knowledge | Fine-tuning adjusts behavior, not knowledge — use RAG |67| No baseline comparison | Always test base model on same eval set |68| Ignoring forgetting | Mix 20% general data to preserve capabilities |