Calibrate Eval Infrastructure

Stop the machine from deciding your benchmark. Configure and validate the container and runtime resources for an agentic coding eval so infrastructure noise stays inside statistical bounds instead of swinging scores more than the models do. Use this whenever someone runs SWE-bench or any agentic coding benchmark in containers, sees scores jump between runs for no code reason, suspects OOM kills or flaky infra are skewing results, sets container memory or CPU limits for an eval harness, or wants to trust a leaderboard delta. Trigger on "my benchmark scores are inconsistent," "OOM during eval," "how much memory should the eval container get," and similar. Not for designing the eval tasks or graders themselves; that's build-agent-evals.

Hoja-Solutions eae24b0 2 files · 7.4 KB Updated

File contents

Hoja-Solutions/agent-stdlib/tree/main/skills/calibrate-eval-infrastructure commit eae24b0485

Frequently asked questions

npx skillmds@latest add hoja-solutions/calibrate-eval-infrastructure