Ds 1000 Eval

This benchmark evaluates a model's ability to generate correct, executable Python code for data science tasks, specifically focusing on NumPy operations. It probes functional correctness under natural language descriptions and tests robustness against surface-form and semantic perturbations of original StackOverflow problems. Use when the user wants to benchmark on numpy-100, or asks about evaluating this task. Reports pass@1.

qhjqhj00 5649927 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/ds-1000-eval commit 56499276e9

Frequently asked questions

npx skillmds add qhjqhj00/ds-1000-eval