Simpo Training

Simple Preference Optimization for LLM alignment. Reference-free alternative to DPO with better performance (+6.4 points on AlpacaEval 2.0). No reference model needed, more efficient than DPO. Use for preference alignment when want simpler, faster training than DPO/PPO.

synthetic-sciences 82602c1 4 files · 31.4 KB Updated

File contents

synthetic-sciences/openscience/tree/main/backend/cli/skills/ml-training/simpo commit 82602c1f99

Frequently asked questions

npx skillmds@latest add synthetic-sciences/simpo-training