Choose Agent Benchmarks

Research, select, audit, reproduce, and design benchmark suites for LLM agents. Use for agent-benchmark surveys, related-work comparisons, evaluation plans, benchmark selection, leaderboard interpretation, reproducibility reviews, or experiments on web agents, computer-use agents, tool/API agents, software-engineering agents, deep-research/scientific agents, memory agents, multi-agent systems, safety/security agents, and embodied agents. Also use when comparing a base model, an agent scaffold/harness, or a complete agent system; when diagnosing saturation, contamination, reward hacking, environment drift, evaluator validity, cost, or reliability; and when studying model–harness co-evolution. Distinguish full interactive benchmarks from static diagnostics, environments, harnesses, and evaluator benchmarks.

pnx2003 Updated

File contents

pnx2003/awesome-agent-skills/tree/main/choose-agent-benchmarks commit b7ca5123a4

Frequently asked questions

npx skillmds@latest add pnx2003/choose-agent-benchmarks