Benchmark Adder

Given a benchmark repository URL (agentic envs like arc-agi, memory/eval benchmarks like longmemeval, QA/code/tool-use benchmarks, etc.), orchestrate the creation of a full Claude Code plugin that benchmarks the current harness setup against it. Wraps the babysitter:babysit skill with the benchmark-plugin-creator process.

tmuskal aab4a00 5.3 KB Updated

File contents

tmuskal/arc-agi-benchmarker/tree/main/.claude/skills/benchmark-adder commit aab4a00eca

Frequently asked questions

npx skillmds@latest add tmuskal/benchmark-adder