Model Evaluation
Principle expression
Primary: P13
Supporting: P02, P05, P08
Scope
Own one recurring judgment: what traceable task evidence justifies a bounded
capability claim about an execution profile strongly enough to inform later
allocation?
Evaluate an execution profile, not a model label. Its identity includes the
model and provider or plan, harness and version, prompt/skill/context policy,
tools and permissions, and execution policy. A changed member creates a new or
revised profile unless evidence shows that the difference is immaterial.
The result is a versioned capability claim with evidence, scope, uncertainty,
and a reopening observation. It is not a universal intelligence score.
Principle source
First detect whether the host declares a Sequence and matching interpretations.
When it does, read only host P13, P02, P05, and P08. Otherwise use the read-only
fallback in references/sequence.md. A live task may select a different current
lead without changing this skill's stable lineage.
Start
Recover only enough state to decide whether a valid evaluation can begin:
Allocation decision the evidence must inform:
Candidate execution profiles and identity revisions:
Real task population and retained acceptance evidence:
Material capability dimensions and failure classes:
Available executor, isolation, usage, latency, and judge evidence:
Existing accepted profile or baseline, if degradation is suspected:
Human or designated host with profile-admission authority:
If there is no allocation decision, real task population, or traceable
acceptance evidence, do not manufacture a benchmark. Return the missing evidence
and the smallest probe that could obtain it.
Comparison validity gate
Apply this gate before interpreting pass rates, latency, judge preference, or
output quality. A later good result cannot repair an invalid comparison design.
- Name the attribution target: the whole execution profile, one deliberately
changed profile member, or a bare model claim. Enumerate every other material
difference between candidates. If several members differ, report only a
whole-profile observation; do not attribute it to the model.
- Compare the worker-visible packet with evaluator-only references. If the
worker can see an expected conclusion, seeded defect, reference answer, or a
semantic criterion that states what it must discover, invalidate capability
comparison. Judge content, not the field label: “inspect every source,”
“cite evidence,” and “produce this artifact shape” are procedural; “X owns
Y,” “A cannot substitute for B,” or “the seeded defect is Z” disclose domain
answers even when placed in an
acceptance field. Source-citation work does
not make a disclosed answer held out. Preserve only protocol or
answer-following evidence.
- Check whether the cases can discriminate the claimed allocation and whether
an execution boundary dominates the outcomes. If all viable profiles can
pass by extraction, or all runs merely hit a cutoff, redesign the field.
State which gate passed or failed in the Capability Case. Continue to outcome
analysis only for the evidence layers that remain valid.
Core method
- Bound the claim. State the task shape, risk, and decision the profile may
influence. “Better model” is not a bounded claim; “reliably reviews local
TypeScript changes before human confirmation” may be.
- Freeze the compared conditions. Record every identity member and keep
task packet, fixture revision, skills, tools, permissions, completion
contract, and execution policy matched unless that member is the variable
under evaluation. Make the declared effective inference policy explicit:
thinking or
reasoning mode and effort, temperature, streaming, context/history handling,
step and duration boundaries, and completion protocol. A selected route
target is not proof of the provider's hidden backend build or revision; name
it as route evidence unless the provider supplies stronger provenance. The
declaration remains a claim until source or provider evidence supports it.
- Select cases from practice. Prefer previously completed production or
project tasks with independent acceptance evidence. Include materially
different cases inside the claimed task population and at least one case
likely to expose a characteristic failure. Public benchmarks may supplement
this field but cannot replace it. Remove cases so easy that every viable
profile passes without exercising the claimed capability, and cases so hard
that every run only hits the execution boundary; neither distinguishes the
allocation decision.
- Name decision-changing evidence before running. For each case state its
acceptance conditions, material failure classes, and the observation that
would defeat the proposed allocation claim. Avoid criteria that reward
verbosity, stylistic resemblance, or test-count inflation. Keep worker-visible
acceptance procedural and artifact-oriented. Expected conclusions, seeded
defects, reference answers, and semantic comparison criteria remain visible
only to the evaluator; otherwise the task measures answer following rather
than discovery.
- Run matched repetitions. Execute each profile more than once in isolated
workspaces. Balance order and avoid parallelism when provider load would add
an uncontrolled variable. Retain unsettled runs, retries, latency, usage,
selected route identity and any stronger provider provenance, artifacts,
verification, and interventions; never
discard a failure to make the sample comparable.
A run stopped at a declared duration boundary is right-censored evidence: it
proves only that the profile did not settle within that envelope. Do not call
one boundary hit random instability or estimate its unseen completion time.
- Judge blind where judgment is necessary. Hide profile/provider/model
identity from the evaluator. Mechanical acceptance may settle deterministic
conditions; semantic judgment reports evidence but cannot admit the profile
as fact. Keep the judge identity and usage visible outside the blind packet.
- Compare variation before averages. Examine within-profile inconsistency,
failure modes, and order effects before claiming a between-profile
difference. If ordinary variance is as large as the claimed advantage,
return
inconclusive and change the next probe rather than adding confidence
language.
- Prepare a Capability Case. Read
references/capability-case.md. Link the
exact fixture, cases, run records, judge evidence, failures, and resource
observations. Treat prompting behavior as evidence about the whole execution
profile: retain observed prompt failures and the smallest treatment hypothesis,
but do not attribute them to the model alone.
- Separate discovery from confirmation. A case used to revise instructions,
skills, context, tool descriptions, or completion contracts has become a
development case. Hold model, route, harness, tools, and fixture fixed while
comparing that prompt treatment; then confirm the revised profile on retained
cases that did not teach the treatment. Do not report tuning-case improvement
as held-out capability.
- Submit deliberately. Only the named human or designated host may accept
the candidate claim into a reusable capability profile. An evaluator,
runtime, provider, or routing policy cannot approve its own allocation.
Execution mechanism
Use the host's smallest truthful executor. It must preserve identity, matched
task input, isolation, repeated outcomes, and raw evidence; the method does not
depend on Work Cell.
When this repository's Work Cell is available, a prepared
work-cell.model-evaluation.v3 manifest can be executed with:
bun packages/work-cell/src/cli.ts model evaluate path/to/model-evaluation.json
The adapter compares exactly two explicit profiles over one declared axis, runs
two to five serial repetitions over a frozen fixture, uses task-local failure
classes and an independent judge route, and emits candidate evidence. Each
profile must name its context, tool-surface, and inference policies. Each case
separates generic worker-visible acceptance from evaluator-only reference
criteria; exact leakage between the two is rejected. Inspect the retained
record; the compact CLI summary intentionally does not name a winner.
When the next question is specifically whether one prompt or skill treatment
improves the same execution member, use the v3 instruction-carrier axis. Hold
the execution fields equal, declare non-empty baseline and treatment carriers,
pin fixture.expectedSha256, and retain a caller-owned semantic-atom audit.
Feed an accepted treatment back as a new profile revision, then return here for
task-population confirmation. Model evaluation discovers prompting hypotheses;
it does not erase the attribution boundary. Prompting variables include wording,
instruction priority and placement, phase separation, tool descriptions, and
completion protocol. Test them separately; do not respond to a protocol failure
by indefinitely adding stronger prose.
Degradation boundary
A surprising answer is not evidence that a model has “become worse.” First
separate model alias, provider routing/load, quota fallback, harness revision,
context delivery, tools, permissions, and transport state.
Only derive a rapid canary after an accepted profile exposes stable high-signal
cases and ordinary variance. Compare current repeated results with that pinned
baseline and report suspected regression only when the change exceeds the
baseline variation. Otherwise return inconclusive. The canary is a projection
of the profile evidence, never an independent benchmark or global degradation
verdict.
Boundaries
- Do not collapse different task dimensions into one score or leaderboard.
- Do not infer capability from price, quota, popularity, or provider marketing.
- Do not let a judge preference erase mechanical failures or divergent runs.
- Do not silently change prompt, skill, context, tools, permissions, or fallback
policy while attributing the result to the model.
- Do not repeatedly tune on evaluation cases and continue calling them held out.
- Do not grant provider or spend authority; route configuration remains an
environment concern.
- Route workload estimation and time/money conversion to their owning method
after a capability claim exists; evaluation records usage and latency only.
- Route one proposed code or artifact comparison to its domain review method.
Completion standard
An evaluation is ready for human submission only when the execution identities,
task population, fixture and acceptance provenance, repeated raw outcomes,
within-profile variance, failure classes, latency and usage, judge identity,
alternative explanation, bounded claim, and reopening observation are explicit.
If any is absent, retain a probe record rather than a capability profile.
1---2name: model-evaluation3description: Build or revise an evidence-linked capability profile for a model execution setup by running repeated, representative real tasks under matched conditions. Use when comparing models, providers, plans, coding harnesses, or prompt/tool profiles for task allocation; when asking "which model is good enough for this work?", "evaluate this model", "compare model capability", "模型评测/能力画像/模型适合什么任务", or whether a characterized setup may have degraded. Do not use for provider setup, public leaderboard summaries, one-off response review, automatic routing, or a degradation verdict without an accepted baseline.4---56# Model Evaluation78## Principle expression910**Primary:** P1311**Supporting:** P02, P05, P081213## Scope1415Own one recurring judgment: **what traceable task evidence justifies a bounded16capability claim about an execution profile strongly enough to inform later17allocation?**1819Evaluate an execution profile, not a model label. Its identity includes the20model and provider or plan, harness and version, prompt/skill/context policy,21tools and permissions, and execution policy. A changed member creates a new or22revised profile unless evidence shows that the difference is immaterial.2324The result is a versioned capability claim with evidence, scope, uncertainty,25and a reopening observation. It is not a universal intelligence score.2627## Principle source2829First detect whether the host declares a Sequence and matching interpretations.30When it does, read only host P13, P02, P05, and P08. Otherwise use the read-only31fallback in `references/sequence.md`. A live task may select a different current32lead without changing this skill's stable lineage.3334## Start3536Recover only enough state to decide whether a valid evaluation can begin:3738```text39Allocation decision the evidence must inform:40Candidate execution profiles and identity revisions:41Real task population and retained acceptance evidence:42Material capability dimensions and failure classes:43Available executor, isolation, usage, latency, and judge evidence:44Existing accepted profile or baseline, if degradation is suspected:45Human or designated host with profile-admission authority:46```4748If there is no allocation decision, real task population, or traceable49acceptance evidence, do not manufacture a benchmark. Return the missing evidence50and the smallest probe that could obtain it.5152## Comparison validity gate5354Apply this gate before interpreting pass rates, latency, judge preference, or55output quality. A later good result cannot repair an invalid comparison design.56571. Name the attribution target: the whole execution profile, one deliberately58 changed profile member, or a bare model claim. Enumerate every other material59 difference between candidates. If several members differ, report only a60 whole-profile observation; do not attribute it to the model.612. Compare the worker-visible packet with evaluator-only references. If the62 worker can see an expected conclusion, seeded defect, reference answer, or a63 semantic criterion that states what it must discover, invalidate capability64 comparison. Judge content, not the field label: “inspect every source,”65 “cite evidence,” and “produce this artifact shape” are procedural; “X owns66 Y,” “A cannot substitute for B,” or “the seeded defect is Z” disclose domain67 answers even when placed in an `acceptance` field. Source-citation work does68 not make a disclosed answer held out. Preserve only protocol or69 answer-following evidence.703. Check whether the cases can discriminate the claimed allocation and whether71 an execution boundary dominates the outcomes. If all viable profiles can72 pass by extraction, or all runs merely hit a cutoff, redesign the field.7374State which gate passed or failed in the Capability Case. Continue to outcome75analysis only for the evidence layers that remain valid.7677## Core method78791. **Bound the claim.** State the task shape, risk, and decision the profile may80 influence. “Better model” is not a bounded claim; “reliably reviews local81 TypeScript changes before human confirmation” may be.822. **Freeze the compared conditions.** Record every identity member and keep83 task packet, fixture revision, skills, tools, permissions, completion84 contract, and execution policy matched unless that member is the variable85 under evaluation. Make the declared effective inference policy explicit:86 thinking or87 reasoning mode and effort, temperature, streaming, context/history handling,88 step and duration boundaries, and completion protocol. A selected route89 target is not proof of the provider's hidden backend build or revision; name90 it as route evidence unless the provider supplies stronger provenance. The91 declaration remains a claim until source or provider evidence supports it.923. **Select cases from practice.** Prefer previously completed production or93 project tasks with independent acceptance evidence. Include materially94 different cases inside the claimed task population and at least one case95 likely to expose a characteristic failure. Public benchmarks may supplement96 this field but cannot replace it. Remove cases so easy that every viable97 profile passes without exercising the claimed capability, and cases so hard98 that every run only hits the execution boundary; neither distinguishes the99 allocation decision.1004. **Name decision-changing evidence before running.** For each case state its101 acceptance conditions, material failure classes, and the observation that102 would defeat the proposed allocation claim. Avoid criteria that reward103 verbosity, stylistic resemblance, or test-count inflation. Keep worker-visible104 acceptance procedural and artifact-oriented. Expected conclusions, seeded105 defects, reference answers, and semantic comparison criteria remain visible106 only to the evaluator; otherwise the task measures answer following rather107 than discovery.1085. **Run matched repetitions.** Execute each profile more than once in isolated109 workspaces. Balance order and avoid parallelism when provider load would add110 an uncontrolled variable. Retain unsettled runs, retries, latency, usage,111 selected route identity and any stronger provider provenance, artifacts,112 verification, and interventions; never113 discard a failure to make the sample comparable.114 A run stopped at a declared duration boundary is right-censored evidence: it115 proves only that the profile did not settle within that envelope. Do not call116 one boundary hit random instability or estimate its unseen completion time.1176. **Judge blind where judgment is necessary.** Hide profile/provider/model118 identity from the evaluator. Mechanical acceptance may settle deterministic119 conditions; semantic judgment reports evidence but cannot admit the profile120 as fact. Keep the judge identity and usage visible outside the blind packet.1217. **Compare variation before averages.** Examine within-profile inconsistency,122 failure modes, and order effects before claiming a between-profile123 difference. If ordinary variance is as large as the claimed advantage,124 return `inconclusive` and change the next probe rather than adding confidence125 language.1268. **Prepare a Capability Case.** Read `references/capability-case.md`. Link the127 exact fixture, cases, run records, judge evidence, failures, and resource128 observations. Treat prompting behavior as evidence about the whole execution129 profile: retain observed prompt failures and the smallest treatment hypothesis,130 but do not attribute them to the model alone.1319. **Separate discovery from confirmation.** A case used to revise instructions,132 skills, context, tool descriptions, or completion contracts has become a133 development case. Hold model, route, harness, tools, and fixture fixed while134 comparing that prompt treatment; then confirm the revised profile on retained135 cases that did not teach the treatment. Do not report tuning-case improvement136 as held-out capability.13710. **Submit deliberately.** Only the named human or designated host may accept138 the candidate claim into a reusable capability profile. An evaluator,139 runtime, provider, or routing policy cannot approve its own allocation.140141## Execution mechanism142143Use the host's smallest truthful executor. It must preserve identity, matched144task input, isolation, repeated outcomes, and raw evidence; the method does not145depend on Work Cell.146147When this repository's Work Cell is available, a prepared148`work-cell.model-evaluation.v3` manifest can be executed with:149150```bash151bun packages/work-cell/src/cli.ts model evaluate path/to/model-evaluation.json152```153154The adapter compares exactly two explicit profiles over one declared axis, runs155two to five serial repetitions over a frozen fixture, uses task-local failure156classes and an independent judge route, and emits candidate evidence. Each157profile must name its context, tool-surface, and inference policies. Each case158separates generic worker-visible acceptance from evaluator-only reference159criteria; exact leakage between the two is rejected. Inspect the retained160record; the compact CLI summary intentionally does not name a winner.161162When the next question is specifically whether one prompt or skill treatment163improves the same execution member, use the v3 `instruction-carrier` axis. Hold164the execution fields equal, declare non-empty baseline and treatment carriers,165pin `fixture.expectedSha256`, and retain a caller-owned semantic-atom audit.166Feed an accepted treatment back as a new profile revision, then return here for167task-population confirmation. Model evaluation discovers prompting hypotheses;168it does not erase the attribution boundary. Prompting variables include wording,169instruction priority and placement, phase separation, tool descriptions, and170completion protocol. Test them separately; do not respond to a protocol failure171by indefinitely adding stronger prose.172173## Degradation boundary174175A surprising answer is not evidence that a model has “become worse.” First176separate model alias, provider routing/load, quota fallback, harness revision,177context delivery, tools, permissions, and transport state.178179Only derive a rapid canary after an accepted profile exposes stable high-signal180cases and ordinary variance. Compare current repeated results with that pinned181baseline and report `suspected regression` only when the change exceeds the182baseline variation. Otherwise return `inconclusive`. The canary is a projection183of the profile evidence, never an independent benchmark or global degradation184verdict.185186## Boundaries187188- Do not collapse different task dimensions into one score or leaderboard.189- Do not infer capability from price, quota, popularity, or provider marketing.190- Do not let a judge preference erase mechanical failures or divergent runs.191- Do not silently change prompt, skill, context, tools, permissions, or fallback192 policy while attributing the result to the model.193- Do not repeatedly tune on evaluation cases and continue calling them held out.194- Do not grant provider or spend authority; route configuration remains an195 environment concern.196- Route workload estimation and time/money conversion to their owning method197 after a capability claim exists; evaluation records usage and latency only.198- Route one proposed code or artifact comparison to its domain review method.199200## Completion standard201202An evaluation is ready for human submission only when the execution identities,203task population, fixture and acceptance provenance, repeated raw outcomes,204within-profile variance, failure classes, latency and usage, judge identity,205alternative explanation, bounded claim, and reopening observation are explicit.206If any is absent, retain a probe record rather than a capability profile.