Agent Benchmark

Use when the user wants a professional, dynamic agent/skill/tool benchmark — compare harnesses, skills, MCPs, CLIs, or workflows on the same tasks with tokens, turns, latency, cost, and success metrics; prove whether a change helps; run ablation-style experiments; or build a reusable bench harness for a repo. Inspired by rigorous same-task evaluation (not GitHub stars).

YosefHayim 475f124 2 files · 9.4 KB Updated

File contents

YosefHayim/dufflebag/tree/main/src/skills/agentBenchmark commit 475f12492f

Frequently asked questions

npx skillmds@latest add yosefhayim/agent-benchmark