Mixeval Eval

Evaluates LLMs on a dynamically mixed benchmark of real-world web-mined queries and existing datasets to measure alignment with human preferences. It probes a model's general capability, reasoning, and instruction-following across diverse domains, correlating performance with Chatbot Arena Elo scores. Use when the user wants to benchmark on MixEval, MixEval-Hard, or asks about evaluating this task. Reports Spearman's ranking correlation.

qhjqhj00 e8ada7d 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mixeval-eval commit e8ada7d0b6

Frequently asked questions

npx skillmds add qhjqhj00/mixeval-eval