Benchmark methodology
Authority: docs/harness/ARCHAEOLOGY.md (the benchmark rows),
benchmarks/METHODOLOGY.md, benchmarks/REPRODUCING.md, CLAUDE.md
Practice 4, docs/harness/FINDINGS.md F-10, F-13, F-17.
Rules, each a pointer, none restated:
- Never hand-type a number. Every figure comes from a run in THIS
session or the CI-captured artifact
benchmarks/jcm_reference.json(Practice 4). The comparators read that file; there is no estimator. - Per row, never per total (F-13): the total hid one cause behind another. Diff threshold ids one by one.
- Five mirrors move together (Practice 4):
results.md,METHODOLOGY.md, README,provenance/measured.json,REPRODUCING.md;tests/test_provenance.pyis the gate and--referencerewrites them. - The reference is captured where the gate runs (F-13):
benchmark.ymldispatch withreference=true; a dev-box number differs by CRLF, walk order and a HOME ledger (F-17). - Deterministic configuration is the bench tier itself:
--offline, pinned corpora (benchmarks/tasks.json), the no-network fixture. Add no flags. - A refusal is not a zero: an absent value prints
n/a. - Floors live only in
harness/thresholds.json; read one withuv run python -m harness threshold <id>. benchmarks/schema_baseline.jsonis the only source for schema-token figures (CLAUDE.md "Tier-switch pricing"); the harness counts a different payload and the two never agree digit for digit.