Codemmlu Eval

Evaluates code understanding and reasoning capabilities of large language models using a multiple-choice question format. It probes syntactic knowledge, semantic comprehension, and real-world software engineering problem-solving, revealing gaps in true code reasoning compared to open-ended generation benchmarks. Use when the user wants to benchmark on CodeMMLU, or asks about evaluating this task. Reports accuracy %.

qhjqhj00 eccef6c 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/codemmlu-eval commit eccef6ccc5

Frequently asked questions

npx skillmds add qhjqhj00/codemmlu-eval