Agentv Bench

Optimize agents through evaluation-driven iteration. Use when asked to evaluate an agent, optimize prompts against evals, run EVAL.yaml or evals.json evaluations, benchmark agent performance, compare agent outputs across providers, analyze eval results, or improve agent performance. Supports workspace evaluation with real repos, multi-provider targets, multi-turn conversations, code judges, tool trajectory scoring, and workspace file change tracking. Use this skill whenever the user mentions evaluating, benchmarking, testing, or optimizing any agent, prompt, or skill — even if they don't explicitly say "agentv".

majiayu000 dc0e39a 2 files · 29.2 KB Updated 567 repo stars

File contents

majiayu000/claude-skill-registry-data/tree/main/testing/agentv-bench commit dc0e39a971

Frequently asked questions

npx skillmds add majiayu000/agentv-bench