Simpo Training

Train language models with SimPO, a reference-free preference optimization method that outperforms DPO without needing a reference model.

Orchestra Research Updated 10.4k repo stars

File contents

Orchestra-Research/AI-Research-SKILLs/tree/main/06-post-training/simpo commit 773a52944b

Frequently asked questions

npx skillmds@latest add orchestra-research/simpo-training