Swe Rebench V2 Eval

Evaluates the ability of LLM-based agents to autonomously resolve software engineering issues by modifying code in real-world repositories. It probes environment setup, code generation, and test execution capabilities across multiple programming languages. Use when the user wants to benchmark on SWE-rebench V2, or asks about evaluating this task. Reports pass@1.

qhjqhj00 98e02c7 2.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/swe-rebench-v2-eval commit 98e02c7bea

Frequently asked questions

npx skillmds add qhjqhj00/swe-rebench-v2-eval