Train Dpo

Direct Preference Optimization (DPO) fine-tune with TRL `DPOTrainer`. Triggered when the user wants to align a model on preferences / pairwise comparisons / chosen-vs-rejected data, or improve an existing SFT checkpoint with a preference dataset.

mybigday Updated

File contents

mybigday/ml-intern-kit/tree/main/.claude/skills/train-dpo commit 19b2ba38da

Frequently asked questions

npx skillmds@latest add mybigday/train-dpo