Multi Swe Bench Eval

This benchmark evaluates an LLM's ability to resolve software engineering issues across multiple programming languages. It probes capabilities in long-context reasoning, multi-file code patching, and fault localization by requiring models to generate executable fixes for real-world repository issues. Use when the user wants to benchmark on Multi-SWE-bench, or asks about evaluating this task. Reports Resolved Rate (%).

qhjqhj00 9ca9b7f 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/multi-swe-bench-eval commit 9ca9b7f974

Frequently asked questions

npx skillmds add qhjqhj00/multi-swe-bench-eval