Task Structure

SWE-bench is the gold-standard benchmark for evaluating AI agents on real-world software engineering tasks.

tools-only Updated 7 repo stars

File contents

tools-only/X-Skills/tree/main/development/1180-swe-bench_b67bf4b4 commit b629a7a2f8

Frequently asked questions

npx skillmds@latest add tools-only/task-structure-14