Open LLM Leaderboard Eval

This benchmark evaluates large language models on their ability to answer open-style questions across diverse knowledge and reasoning domains. It specifically probes whether models rely on selection bias and random guessing in multiple-choice formats versus genuinely understanding and generating correct open-ended responses. Use when the user wants to benchmark on MMLU, ARC, MedMCQA, PIQA, CommonsenseQA, OpenBookQA, RACE, WinoGrande, or asks about evaluating this task. Reports accuracy.

qhjqhj00 cdbaf06 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/open-llm-leaderboard-eval commit cdbaf06030

Frequently asked questions

npx skillmds add qhjqhj00/open-llm-leaderboard-eval