Livebench Eval

Evaluates large language models across 18 tasks spanning math, coding, reasoning, language, instruction following, and data analysis. It uses dynamically updated, objectively scored questions from recent real-world sources to minimize test-set contamination and avoid LLM-judging biases. Use when the user wants to benchmark on LiveBench, or asks about evaluating this task. Reports LiveBench score.

qhjqhj00 88447d2 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/livebench-eval commit 88447d2df6

Frequently asked questions

npx skillmds add qhjqhj00/livebench-eval