Benchmark Harness

Evaluate the coding harness you are running inside (pi, opencode, claude-code) on its live setup and current model, via bench setup run. Measures it with and without its skills, MCP servers and plugins on --ab, and reports which of them the run actually called. Use when the user runs /benchmark-harness. Not for serving sweeps, model compare, or one-shot suites.

luongnv89 6e01dd8 11 files · 76.4 KB Updated

File contents

luongnv89/m-bench/tree/main/.agents/skills/benchmark-harness commit 6e01dd8701

Frequently asked questions

npx skillmds@latest add luongnv89/benchmark-harness