Dpo Ppo Multi Bench Eval

This protocol evaluates language models across factual knowledge, mathematical reasoning, instruction following, code generation, truthfulness, and safety/refusal capabilities. It uses standardized benchmarks to measure how preference optimization methods and data quality impact model performance. Use when the user wants to benchmark on MMLU, GSM8k, Big Bench Hard, TruthfulQA, AlpacaEval, IFEval, HumanEval+, MBPP+, ToxiGen, XSTest, or asks about evaluating this task. Reports average accuracy.

qhjqhj00 5fc24d4 5.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/dpo-ppo-multi-bench-eval commit 5fc24d4d21

Frequently asked questions

npx skillmds add qhjqhj00/dpo-ppo-multi-bench-eval