Preference Optimization

Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO. Use when preference pairs or thumbs-up/down feedback exist, when choosing between preference-optimization methods, or when a DPO run needs hyperparameters or debugging.

wshobson 4f3af04 2 files · 14.3 KB Updated

File contents

wshobson/agents commit 4f3af04cca

Frequently asked questions

npx skillmds add wshobson/preference-optimization