Omni Dpo Eval

Evaluates the instruction-following and mathematical reasoning capabilities of LLMs fine-tuned with a dual-perspective preference optimization method. It measures conversational quality, adherence to instructions, and problem-solving accuracy across diverse open-ended and quantitative benchmarks. Use when the user wants to benchmark on AlpacaEval 2.0, Arena-Hard v0.1, IFEval, SedarEval, GSM8K, MATH 500, AIME 2024, AMC 2023, or asks about evaluating this task. Reports LC(%).

qhjqhj00 f574c0f 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/omni-dpo-eval commit f574c0f615

Frequently asked questions

npx skillmds add qhjqhj00/omni-dpo-eval