Alpaca Eval Lc Winrate Eval

Evaluates the alignment quality of language models by measuring their win rate against a baseline on the AlpacaEval benchmark. It specifically uses length-controlled (LC) win rates to mitigate the known bias toward longer model outputs in standard auto-annotator evaluations. Use when the user wants to benchmark on alpaca_eval, or asks about evaluating this task. Reports AlpacaEval length-controlled (LC) win rate.

qhjqhj00 eb6cc59 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/alpaca-eval-lc-winrate-eval commit eb6cc59464

Frequently asked questions

npx skillmds add qhjqhj00/alpaca-eval-lc-winrate-eval