Chatbot Arena Eval

Evaluates large language models by collecting human preference votes on pairwise responses to real-world prompts, then ranks them using Bradley-Terry models to measure alignment and real-world utility. Use when the user wants to benchmark on Chatbot Arena, or asks about evaluating this task. Reports BT coefficients.

qhjqhj00 1458132 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/chatbot-arena-eval commit 1458132935

Frequently asked questions

npx skillmds add qhjqhj00/chatbot-arena-eval