Dpo Preference Eval

This evaluation protocol assesses a language model's ability to align with human preferences across open-ended text generation tasks. It measures how well the model optimizes a reward objective while staying close to a reference policy, and evaluates practical performance via pairwise win rates against baselines. Use when the user wants to benchmark on IMDb, Reddit TL;DR, Anthropic HH, or asks about evaluating this task. Reports win rate.

qhjqhj00 0c1ac0e 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/dpo-preference-eval commit 0c1ac0e37f

Frequently asked questions

npx skillmds add qhjqhj00/dpo-preference-eval