Harness Benchmarking

Use when designing evaluation suites for agentic systems, choosing agent benchmarks, measuring multi-agent performance, or critiquing single-score evaluation of agents. Based on arXiv:2605.26112v1. Trigger phrases: "evaluate my agent", "benchmark agent", "measure agent reliability", "pass^k evaluation", "agent metrics beyond success rate", "trajectory quality", "memory hygiene metric", "longitudinal agent eval", "why my agent scores well on benchmarks but fails in production".

mouadja02 bc93087 17.7 KB Updated

File contents

mouadja02/skills/tree/main/skills/agent-eval/harness-benchmarking commit bc9308725b

Frequently asked questions

npx skillmds@latest add mouadja02/harness-benchmarking