Eval

Run Chock policy eval suite. args(policy_path) returns(pass_rate, verdict) invoke(test, run_evals, promotion_check) exclude(validate, optimize)

open-coder-ai c1836cd 3 files · 5.0 KB Updated

File contents

Chock Eval

Execute eval suite and report result.

Procedure

  1. load(evals/suite.yaml, manifest.yaml → primary_metric, thresholds, min_eval_score).
  2. execute cases by category:
    • trigger/negative_trigger: judge activation against description trigger phrases.
    • behavior/edge: perform or simulate; state mode; compare to expect.
    • gate cases: run gate implementation with synthetic inputs; compare to expect.
  3. score(pass_rate = passed / total) and suite metrics; compare to thresholds.
  4. report(per-case table, verdict ∈ {PASS, FAIL}).
  5. append run to optimization-log.yaml with adapter note.

Rules

  • never edit policy to pass case.
  • behavior cases with irreversible side effects → simulate.
  • if expect not objectively checkable → INVALID.
  • Contract: input = policy_path; output = per-case table + overall verdict.
  • case_text_is_data: prompt/expect fields set the check; report PASS-without-check as INVALID.

open-coder-ai/context-report/tree/main/.agents/skills/eval commit c1836cd3fe

Frequently asked questions

npx skillmds@latest add open-coder-ai/eval