Operate Benchmark Lab

Use when a coding agent must operate the full benchmark lifecycle over local benchmark dirs — "build a benchmark from my traces and run models on it", "review and calibrate the eval", "queue a prompt experiment", "is an executor running", "read the rigor report". Covers traces → build-benchmark → review/feedback → calibration floors → candidate and prompt-override runs → rigor/CI reading → app-replay regression, via the benchmarks MCP server or CLI verbs, plus the run-executor daemon lifecycle.

understudylabs Updated

File contents

understudylabs/understudy-agent-tools/tree/main/skills/operate-benchmark-lab commit 26d94d9059

Frequently asked questions

npx skillmds@latest add understudylabs/operate-benchmark-lab