Bigcodebench Eval

Evaluates large language models' ability to generate correct, executable code for complex programming tasks requiring diverse function calls and compositional reasoning. It also probes instruction-following capabilities by comparing performance on verbose prompts versus condensed natural-language instructions. Use when the user wants to benchmark on BigCodeBench, or asks about evaluating this task. Reports Pass@1.

qhjqhj00 31a7af8 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/bigcodebench-eval commit 31a7af8f2c

Frequently asked questions

npx skillmds add qhjqhj00/bigcodebench-eval