Swe Bench Java Eval

This benchmark evaluates an AI agent's ability to autonomously resolve real-world GitHub issues in Java projects. It probes capabilities in code patch generation, repository navigation, test case reasoning, and handling runtime environment dependencies. Use when the user wants to benchmark on SWE-bench-java-verified, or asks about evaluating this task. Reports Resolved Rate (%).

qhjqhj00 ecab78e 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/swe-bench-java-eval commit ecab78efea

Frequently asked questions

npx skillmds add qhjqhj00/swe-bench-java-eval