Agent Arena

Use when the user asks for a second opinion, independent/heterogeneous review, architecture red-team, cross-model critique (Codex/Claude/GLM/DeepSeek/Qwen/Kimi), review my plan, challenge this design, evidence-checked code/PR review, or multi-agent critique of a high-stakes plan, design, research claim, or bug root-cause. Not for simple lookups, formatting, or low-stakes tasks. On error_max_turns before any answer: mechanical failure, not a result — check if you boxed an open design review (needs broad read-only tools + ample turns) as bounded verification; retry ONCE with lossless moves only. Disabling the reviewer's tools, feeding excerpts, or narrowing scope are LOSSY — never automatic; STOP and ask the user. Keep packet AND output on disk (stage packet via stdin, redirect claude -p output to a file); read back only a structured digest preserving dissent, never the raw JSON; checkpoint each round to disk — else the prompt+output inflate your context, trigger compaction, and loop into re-running arena.

zhjai Updated

File contents

zhjai/agent-arena/tree/main/skills/agent-arena commit b028522de1

Frequently asked questions

npx skillmds@latest add zhjai/agent-arena