Cruxevalx Eval

This benchmark evaluates large language models' ability to perform bidirectional code reasoning across 19 programming languages. It probes whether models can predict missing inputs given outputs, and predict missing outputs given inputs, testing their understanding of code semantics and execution flow. Use when the user wants to benchmark on CRUXEval-X, or asks about evaluating this task. Reports Pass@1.

qhjqhj00 9a103eb 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/cruxevalx-eval commit 9a103ebeef

Frequently asked questions

npx skillmds add qhjqhj00/cruxevalx-eval