Agentv Bench

Run AgentV evaluations and optimize agents through eval-driven iteration. Triggers: run evals, benchmark agents, optimize prompts/skills against evals, compare agent outputs across providers, analyze eval results, offline evaluation of recorded sessions, run autoresearch, optimize unattended, run overnight optimization loop. Not for: writing/editing eval YAML without running (use agentv-eval-writer), analyzing existing traces/JSONL without re-running (use agentv-trace-analyst).

EntityProcess b2b2b5a 15 files · 154.4 KB Updated

File contents

EntityProcess/agentv/tree/main/skills-data/agentv-bench commit b2b2b5a003

Frequently asked questions

npx skillmds@latest add entityprocess/agentv-bench