Livecodebench Eval

Evaluates large language models' ability to generate, repair, execute, and predict outputs for code across multiple algorithmic and competitive programming problem types. It provides a holistic, contamination-free assessment by continuously updating with new problems and testing models across four distinct coding scenarios. Use when the user wants to benchmark on LiveCodeBench, or asks about evaluating this task. Reports PASS@1.

qhjqhj00 7f48f3e 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/livecodebench-eval commit 7f48f3e1ec

Frequently asked questions

npx skillmds add qhjqhj00/livecodebench-eval