Codebase Arena

Build, execute, and judge human-reviewed Chinese-English bilingual codebase capability benchmarks. Use when an agent must generate a repository-specific evalset, run an approved evalset inside a tested desktop coding agent through fresh single-case subagents, or independently judge and compare completed raw runs. Route to the generation, execution, or report workflow; require structured human intake at the start of every new workflow; keep private result-verification material unavailable to evaluated agents; generate only read-only code-understanding cases with no code-generation load while preserving legacy execution/report compatibility; produce one 0-10 score per system and case; and never install or change dependencies without explicit user approval.

codeartsagent 4be4134 47 files · 461.6 KB Updated

File contents

codeartsagent/codeartsskills/tree/main/skills/codebase-arena commit 4be41340eb

Frequently asked questions

npx skillmds@latest add codeartsagent/codebase-arena