Swe Chat Eval

Evaluates real-world coding agent interactions by measuring how much agent-generated code survives into final commits, alongside efficiency metrics like token usage, cost, and runtime per committed line. Use when the user wants to benchmark on SWE-chat, or asks about evaluating this task. Reports Code survival rate.

qhjqhj00 240a807 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/swe-chat-eval commit 240a80763d

Frequently asked questions

npx skillmds add qhjqhj00/swe-chat-eval