Alpacaeval Eval

Evaluates LLM response quality via pairwise win rates against a baseline, while specifically probing the metric's susceptibility to length bias, gameability via verbosity prompting, and robustness to adversarial truncation. Use when the user wants to benchmark on AlpacaEval, or asks about evaluating this task. Reports Win rate.

qhjqhj00 9957563 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/alpacaeval-eval commit 9957563ba9

Frequently asked questions

npx skillmds add qhjqhj00/alpacaeval-eval