Finetuning

Use when adapting an open-weight model to a target form or behavior — tone, output format, reasoning pattern — via LoRA/QLoRA or full fine-tuning with TRL SFTTrainer, then preference optimization (DPO/ORPO/KTO/GRPO), and for fine-tune vs prompt vs RAG. NOT adding facts to a model (that is rag); NOT the single-GPU Unsloth backend or GGUF export (that is unsloth).

ericrisco c6d4e4b 5 files · 33.2 KB Updated

File contents

ericrisco/rsc-harness/tree/main/skills/finetuning commit c6d4e4be94

Frequently asked questions

npx skillmds@latest add ericrisco/finetuning