Babelcode Eval

Evaluates the capability of large language models to generate executable code and translate code across multiple programming languages. It measures functional correctness by executing generated programs against test cases and computing the probability that at least one sample passes. Use when the user wants to benchmark on BC-HumanEval, BC-MBPP, BC-Transcoder, TP3, or asks about evaluating this task. Reports pass@k.

qhjqhj00 61a7206 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/babelcode-eval commit 61a7206060

Frequently asked questions

npx skillmds add qhjqhj00/babelcode-eval