Benchmark Operator Skill
Use this skill when the task is benchmark orchestration instead of normal governed delivery work.
Workflow
- Select the suite and subjects.
- Run the suite or import manual outputs.
- Review catch rate, block rate, expected-rule hit rate, latency, and overhead.
- Export the run if you need external analysis.
Local observability feedback loop
For generated observability coverage, prefer the human-gated loop:
- Inspect
ck_observabilityreports forsaved_evals,benchmark_drafts,benchmark_scenarios, andbenchmark_history. - Use CLI-only commands for mutations: draft approval, materialization, and benchmark execution are not exposed through MCP in this skill.
- Run generated observability benchmarks only after a dry-run review and explicit operator approval.
- Treat
promotionsas advisory evidence; do not mutate policy, router, prompt, or autofix artifacts from benchmark results alone.
Additional resources
- Benchmark operator playbook