Mceval Eval

Evaluates the multilingual code generation, explanation, and completion capabilities of LLMs across 40 programming languages. It measures how well models can produce correct code, explain code logic, and complete code snippets in diverse syntaxes. The benchmark highlights performance disparities between closed-source and open-source models, particularly in non-Python languages. Use when the user wants to benchmark on MCEVAL, or asks about evaluating this task. Reports Pass@1 (%).

qhjqhj00 9ac7a43 2.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mceval-eval commit 9ac7a43fbe

Frequently asked questions

npx skillmds add qhjqhj00/mceval-eval