Eval Gate Authoring
Intro
Eval gate authoring promotes repeated run observations into enforceable
checks. Capture examples, codify an eval-spec Artifact, create the
paired Gate, calibrate any LLM judge against human labels, then bind
the eval to the target run surface with a policy-application Binding.
Overview
The workflow starts from observed run output and ends with an auditable Gate plus bindings that describe where the evaluation applies. LLM-as-judge evals require calibration evidence before they should be treated as enforceable.
Gotchas
- Do not create an eval gate from one example unless the gate is clearly marked exploratory.
- Do not treat an LLM judge as calibrated until human labels and a calibration LogEntry exist.
- Bind evals to the narrowest useful run surface so a local experiment does not accidentally become a global policy.
Full reference
MCP tools
collect_run_outputscodify_evalcalibrate_judgebind_eval_to_runs
Produced entities
Artifact(spec.kind=eval-spec)GateBinding(type=policy-application)- calibration
LogEntryrecords