Swe Bench Repair Eval

This evaluation probes an LLM-based agent's ability to automatically locate faults and generate correct code patches for real-world software issues. It tests both traditional text-only bug fixing and multimodal reasoning where visual UI behavior must be understood alongside code. Use when the user wants to benchmark on SWE-bench Lite, SWE-bench Multimodal, or asks about evaluating this task. Reports %Resolved.

qhjqhj00 3caf4ae 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/swe-bench-repair-eval commit 3caf4ae7fd

Frequently asked questions

npx skillmds add qhjqhj00/swe-bench-repair-eval