Battle

Benchmark AI model combinations on the same coding task. Runs Claude, Antigravity CLI (agy — replaces the sunset Gemini CLI on consumer plans as of 2026-06-18), Codex CLI, Kimi (Moonshot API), and DeepSeek (OpenAI-compatible API) solo and in pairs, then scores each on tokens, cost, wall time, and output quality. Produces a leaderboard to inform future model selection. Use when asked to "battle", "benchmark models", "compare models", "which model is best", "pair programming battle", or "/battle".

eprouveze 7b9dad1 9 files · 28.3 KB Updated

File contents

eprouveze/claude-skills/tree/main/skills/battle commit 7b9dad1809

Frequently asked questions

npx skillmds@latest add eprouveze/battle