Csr Bench Eval

Evaluates LLM agents' ability to autonomously deploy computer science research repositories by executing multi-stage workflows including setup, downloading dependencies, training models, running evaluations, and performing inference. It probes instruction comprehension, command generation, and iterative error correction in complex software environments. Use when the user wants to benchmark on CSR-Bench, or asks about evaluating this task. Reports success rate.

qhjqhj00 39f2528 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/csr-bench-eval commit 39f2528e76

Frequently asked questions

npx skillmds add qhjqhj00/csr-bench-eval