PM AI Independent Evaluation to Release
Turn an external AI evaluation into a reviewable product decision. The unit of
work is not a headline score; it is the chain from claim to tested system,
harness, budget, validity checks, evidence, limitations, remediation, and
deployment action.
Independent evaluation can add perspective alongside internal testing, but it
does not become a safety certificate, compliance opinion, adoption signal, or
truth judgment. Keep the evaluator's result, the product's controls, and the
user's outcome as separate evidence layers.
When to use
Use this skill when:
- a PM is commissioning or reviewing an independent lab, external expert, or
third-party assessment before deploying an AI model, agent, or product;
- an evaluation needs to support a defined capability, safeguard, or system
comparison claim rather than an unspecified “quality” claim;
- an agentic result may depend on tools, memory, retries, scaffolding, context
management, environment state, or a product-facing harness;
- an external report needs a methodology review for scope, independence,
conflicts, budget, task validity, contamination, refusals, reward hacking,
evaluation awareness, or broken environments;
- a team must define safe model/checkpoint access, data retention, credentials,
confidentiality, responsible disclosure, redaction, or publication terms;
- the result may change
Ship, Pilot, Hold, Need evidence, or Rollback
for a consequential AI capability.
Use pm-ai-evaluation-plan when the primary job is designing a general test
plan before an evaluator or result exists. Use
pm-ai-review-to-calibration when the primary job is aligning human or model
judgments. Use pm-ai-prompt-injection-to-defense when the primary job is one
untrusted-content attack path and its product controls. Use
pm-ai-risk-to-control for a broad hazard register, and
pm-ai-incident-to-runbook after a journey-level incident has occurred.
Do not use this when
- the request is to run a live jailbreak, exploit, attack, scanner, or provider
call;
- the only input is a vendor benchmark and there is no product claim or tested
system to bound;
- the team wants a generic score, a public ranking, a certification, or a
promise of production readiness;
- a red-team exercise is being treated as the whole evaluation; red teaming is
a risk-discovery method and must be scoped separately from capability,
methodology, and product evaluation;
- the open decision is primarily user approval, authorization, moderation,
provenance, model migration, cost, or user retention;
- private customer data, hidden prompts, credentials, exploit payloads, or
chain-of-thought would need to be copied into the public packet.
Evidence boundary
Record each statement as Observed, Reproduced, Supplied, Proposed,
Not run, Unknown, or Not covered. A source may describe a method without
proving that the method was applied to this system.
| Evidence layer |
It may establish |
It cannot establish by itself |
provider_or_standard |
documented evaluation, harness, or safeguard capability |
this product's result, adoption, or deployment fitness |
evaluator_process |
who tested, under what independence, scope, access, and method |
that the evaluator had representative coverage or no bias |
tested_configuration |
model/version, prompt, tools, memory, environment, budget, and safeguards actually in scope |
behavior after a configuration, harness, or policy change |
evaluation_result |
observed scores, examples, failures, and analysis under the declared setup |
a capability ceiling, real-world prevalence, or generalization beyond scope |
product_control |
deterministic gate, approval, sandbox, monitor, or fallback behavior |
model capability or user outcome |
deployment_outcome |
a verified product run, pilot observation, or user decision at a named surface |
broad adoption, safety, or causality without its own evidence |
If a report omits system identity, harness, budget, denominator, evaluator
conflict, task validity, or limitation, lower the claim or return Need evidence. Never turn a missing field into a passing assumption.
Core definitions
| Term |
Working meaning |
Do not confuse it with |
Evaluation claim |
The precise capability, safeguard, or comparison statement the assessment is intended to support |
a broad “the model is good” statement |
Independent evaluation |
An external assessment whose methods and conclusions are not fully controlled by the product team |
a vendor marketing page or internal QA run |
Methodology review |
External scrutiny of the framework, data, scoring, or interpretation |
a rerun of the product itself |
SME probing |
Domain-expert testing of real-world tasks with structured input |
representative user research or adoption evidence |
Red teaming |
Proactive exploration of risks or attacks to discover issues and build evaluations |
a complete deployment evaluation |
Harness |
Model-facing prompts, tools, interfaces, memory, retries, validators, control logic, and environment |
the model name alone |
Maximum elicitation |
Testing for the strongest credible behavior under a defined effort/budget |
an unlimited or undisclosed attempt count |
Validity hazard |
A condition that can inflate or suppress the measured result, such as reward hacking, refusal, contamination, broken task, or evaluation awareness |
ordinary variance |
Evidence packet |
Versioned claim, configuration, run, artifacts, limitations, and decision records |
a slide with one score |
Publication boundary |
What can be shared, redacted, delayed, corrected, or withheld, with an owner |
permission to publish every raw artifact |
Release decision |
Ship, Pilot, Hold, Need evidence, or Rollback with conditions and owner |
safety, legal, compliance, or adoption certification |
Workflow
1. Frame the claim and decision
Write one sentence:
Decide whether evidence from <evaluation> supports <specific claim> for
<user/job or deployment decision> under <scope, risk, and date>.
State the current workaround, audience, decision owner, false-positive and
false-negative costs, version boundary, and what action the result could change.
Choose one primary claim class:
- Capability elicitation: what the system can plausibly produce under a
credible setup;
- Safeguard performance: how robust a defined control is against a defined
behavior or adversary model;
- Controlled comparison: how systems perform under equivalent conditions;
- Methodology / assurance: whether a test design or interpretation is fit
for the decision;
- SME probing: what qualified experts observe on domain tasks.
Do not combine these into a single score. If the claim is “safe in production,”
split it into model behavior, product controls, user experience, operations,
and deployment evidence.
2. Establish independence, access, and authority
Create an evaluator ledger before accepting results:
| Field |
Required question |
Status |
evaluator |
Who performed, funded, supervised, and reviewed the work? |
Observed / Unknown |
expertise |
Which capability or risk area is in scope, and what qualification supports it? |
Observed / Not provided |
conflicts |
What financial, publication, access, or relationship conflicts exist? |
Observed / Unknown |
scope |
Which model, checkpoint, product surface, locale, tools, policies, and dates were tested? |
Observed / Not provided |
access |
Which credentials, data, retention, logging, chain-of-thought, or reduced-safeguard access was granted? |
Observed / Not provided |
authority |
Who can approve test changes, disclose a finding, remediate, or decide release? |
Observed / Not provided |
publication |
Which raw artifacts, summaries, limitations, and corrections can be shared? |
Observed / Not provided |
Treat evaluator access as least privilege. Do not place real customer data,
production secrets, private model weights, or unrestricted external write
authority into a test environment unless the access decision, retention,
isolation, owner, expiry, and rollback are explicit.
3. Freeze the tested system and harness
Assign stable IDs for the evaluation, model/provider/version, prompt or policy
version, tool and memory versions, task/environment snapshot, harness, budget,
scorer, reviewer, and report. For an agent, record:
- interface and user-visible surface, not only a raw model endpoint;
- tools, schemas, retries, timeouts, context/compaction, memory, validators,
approvals, sandbox, and egress boundaries;
- initial state, external data, seed or randomization, locale, and policy;
- turns, tokens, attempts/retries, wall-clock budget, inference cost, and
expected cost per successful solve where relevant;
- model-level and product-level mitigations that were enabled, disabled, or
changed during the run.
If a report used a simpler harness than the real product, state the resulting
claim as a component or lower-bound observation. If the harness changed between
systems or runs, label the comparison Not comparable until the difference is
explained or the evaluation is repeated.
4. Define tasks, slices, oracle, and denominator
Build a small, versioned evaluation matrix before looking at the headline:
| ID |
Claim slice |
Task / fixture |
Expected behavior |
Oracle |
Hard blocker |
Status |
E-001 |
core user job |
safe fictional task |
observable completion |
deterministic or human reference |
none |
Not run |
E-002 |
negative route |
should abstain or ask |
no unsupported assertion/action |
route or policy oracle |
unsafe action |
Not run |
E-003 |
edge / OOD |
locale, format, or missing tool |
bounded recovery |
trace plus review |
silent failure |
Not run |
E-004 |
validity |
scorer or environment trap |
no shortcut credit |
sample inspection |
reward hack |
Not run |
E-005 |
safeguard |
safe category fixture |
control blocks or escalates |
product control receipt |
privacy/security/action leak |
Not run |
Use private or newly constructed tasks when contamination matters. Keep golden,
regression, negative, high-risk, OOD, and holdout slices distinct. Report
counts, missing cases, abstentions, retry/attempt policy, and the denominator.
A rate without its unit, sample, or budget is not a complete result.
5. Check validity hazards before interpreting results
For every result, record checked, not checked, or not applicable for:
- Reward hacking: the system earns credit through a shortcut in the task,
scorer, prompt, tool, or harness rather than the intended behavior;
- Refusals: safeguards suppress the tested behavior and make capability or
robustness difficult to interpret;
- Contamination: tasks, answers, or close variants were in training data or
discoverable through tools;
- Broken problems: ambiguous ground truth, missing files, flaky services,
impossible environments, or unfair scoring;
- Evaluation awareness / sandbagging: the system changes behavior because
it recognizes evaluation or the consequences of a high score;
- Harness and budget drift: retries, tools, context preservation, scoring,
or effort differ in ways that change the observed result;
- Review drift: a judge, rubric, policy, or human interpretation changed
after seeing the result.
Inspect representative trajectories, not only summary scores. A higher score
with more shortcuts, malformed tool structure, or hidden refusal behavior is
not automatically an improvement.
6. Separate result, limitation, and product control
Write findings in this order:
- Observation: what happened under the frozen setup.
- Interpretation: what the result may support and what it cannot support.
- Limitation: which task, model, harness, budget, validity, or access gap
narrows the claim.
- Product consequence: what control, UX state, human review, or scope
restriction should change.
- Re-test: which artifact or condition must be rerun after remediation.
Keep a model refusal separate from a product control that prevented a tool call,
blocked an egress path, or gave a user a safe recovery. Keep a capability result
separate from a user outcome. An external evaluator can provide additional
evidence; it does not own the product's final release decision.
7. Set data, confidentiality, and publication boundaries
For each artifact, define:
| Artifact |
Minimum public form |
Restricted material |
Owner / expiry |
| task set |
category, count, sampling rule, safe fixture |
private prompts, live credentials, sensitive targets |
Not provided |
| trace / trajectory |
redacted step summary or stable ID |
secrets, PII, hidden reasoning, private URLs |
Not provided |
| score / report |
claim, setup, denominator, limitations, result |
confidential details, exploit-enabling material |
Not provided |
| finding |
sanitized issue, severity, mitigation, retest state |
weaponizable payload or private incident detail |
Not provided |
Record data retention, deletion/correction, access logging, responsible
disclosure, accuracy review, publication delay, and correction path. If a result
cannot be independently inspected because of confidentiality, narrow the claim
and label the evidence boundary instead of implying full transparency.
8. Choose the smallest release decision
Ship: the claim is narrow and supported by the correct harness, critical
validity checks are addressed, no hard blocker is open, ownership and rollback
are real, and the product controls were tested at the surface that will ship;
Pilot: exposure is narrow, high-impact actions are gated, human review or
safe fallback exists, unresolved limitations have owners and dates, and the
pilot has an observable stop rule;
Hold: evaluator independence, configuration, task validity, denominator,
access, publication, or control evidence is missing or not comparable;
Need evidence: the decision is well framed but the relevant evaluation
has not run or the supplied result cannot be verified;
Rollback: a released claim or control is contradicted by a verified
critical finding, regression, access failure, or unsafe product behavior.
For every choice, name the user-visible state, owner, TTL/review date, fallback,
rollback trigger, and one next learning question. Restore requires a new
verification window; silence is not evidence of recovery.
9. Write back and ask one review question
Return a packet that an evaluator, PM, engineering owner, safety/security
reviewer, and release owner can all inspect. Link the report, configuration,
validity notes, findings, remediation, re-test, publication status, and decision
without copying restricted content. Add a new regression, negative case, or
methodology correction to the authorized evaluation record when appropriate.
End with one concrete ask, such as:
Approve the narrow pilot with the stated stop rule.
Hold until the harness and denominator are frozen.
Add the validity failure to the holdout and rerun.
Need the evaluator conflict, access, or publication record.
Output contract
Return these sections in order:
## Evaluation decision on the desk
## Claim and user/job boundary
## Evaluator independence and access ledger
## Tested system and harness
## Task, slice, oracle, and denominator matrix
## Validity hazard review
## Result and limitation ledger
## Product control and remediation
## Data, confidentiality, and publication boundary
## Release, fallback, and rollback
## Not covered
## Review ask
Use Not provided, Unknown, Not run, Not comparable, Not verified,
Proposed, or Not covered when the evidence does not support a stronger
statement. For a copy-ready fictional shape, read
references/independent-evaluation-release-brief.md.
Common rationalizations and red flags
| Rationalization |
Red flag / correction |
| “The external lab is credible, so the score speaks for itself.” |
Credibility does not replace a claim, configuration, harness, denominator, or validity review. |
| “It is the same model name, so the comparison is fair.” |
Tools, retries, context, policy, budget, and product surface can change the result. Freeze or label the comparison. |
| “The benchmark is public and large, so it proves generalization.” |
Public or reused tasks can be contaminated; report the tested distribution and limitation. |
| “The model refused, so the product is safe.” |
Inspect tool calls, egress, data boundaries, user recovery, and product controls separately. |
| “Red teaming found no issue, so we can ship.” |
Red teaming is one risk-discovery slice, not a complete capability, product, or deployment evaluation. |
| “We can publish the raw trace to build trust.” |
Redact secrets, PII, hidden instructions, private URLs, and exploit-enabling detail; publish the smallest useful artifact. |
| “The report is confidential, so limitations do not need to be stated.” |
Confidentiality narrows the claim; it does not justify silent uncertainty. |
| “A green CI run makes the external assessment complete.” |
CI proves repository checks, not evaluator independence, harness validity, product behavior, or adoption. |
Edge cases
- No external evaluator: produce a commissioning brief with
Result: Not run and keep the decision at Need evidence.
- Vendor-produced benchmark: record it as
Supplied, name the vendor's
interest and setup, and do not label it independent.
- Model-only test for an agent product: state the product harness gap and
do not generalize the model result to the agent.
- Different harnesses: report separate results or rerun under a shared
setup; do not rank systems from an unqualified comparison.
- Reduced safeguards or early checkpoint: record the exact access boundary,
isolation, retention, and the fact that the result may not represent the
deployed configuration.
- High score with reward hacking or broken tasks: quarantine the result,
correct the task/scorer, and return
Hold or Need evidence.
- High-risk red-team finding: keep the safe summary public, restrict the
payload, assign remediation, and do not claim the absence of a finding means
the system is secure.
- Small sample or no denominator: keep the result descriptive and label the
decision
Not measurable or Need evidence.
- Confidential third-party report: preserve the report locator, scope,
methodology review, and limitation statement without exposing restricted text.
- Fictional or synthetic fixture: label it clearly; it supports workflow
shape and known edge cases, not real capability, safety, user, or adoption.
Final check
Before returning the packet, confirm:
- the decision, user/job, claim class, evaluator, system, version, harness,
budget, task set, denominator, and observation window are visible or marked
missing;
- independence, conflict, access, data retention, confidentiality,
publication, and responsible-disclosure boundaries are explicit;
- capability, safeguard, comparison, methodology, product control, and user
outcome evidence are not collapsed into one score;
- golden, negative, edge/OOD, validity, high-risk, and holdout slices are
represented where relevant;
- reward hacking, refusals, contamination, broken problems, evaluation
awareness, harness/budget drift, and review drift are checked or marked;
- product controls, human review, fallback, rollback, owner, TTL, and stop
condition are observable;
- the report distinguishes supported, unsupported, unknown, not run, blocked,
and not covered claims;
- no secret, private payload, raw customer content, hidden reasoning, or
unsupported safety, truth, legal, adoption, or production claim was added.
References and first run
- For a copy-ready fictional workflow, read
examples/first-run.md. It is a fictional
fixture, not a live evaluation or deployment result.
- For a complete packet shape and source ledger, read
references/independent-evaluation-release-brief.md.
- For current technical source context, read the official sources in the
reference file; refresh them before making version-sensitive claims.
Not covered
- This skill does not run an evaluation, red team, jailbreak, exploit, scanner,
model call, external provider, or third-party engagement.
- It does not certify safety, security, compliance, legal status, truth,
authorship, production readiness, user satisfaction, adoption, or GitHub
growth.
- It does not authorize access to private data, model weights, credentials,
hidden reasoning, production tools, or external write surfaces.
1---2name: pm-ai-independent-eval-to-release3description: Turn an independent or third-party AI evaluation into an evidence-bounded release decision covering claim type, evaluator independence, system and harness configuration, budget, validity hazards, access and publication boundaries, remediation, and rollback. Use when a PM needs to commission, review, or interpret an external evaluation of an AI model, agent, capability, safeguard, or product before deployment; do not treat one report, benchmark score, red-team exercise, or provider statement as proof of safety, truth, adoption, or production readiness.4---56# PM AI Independent Evaluation to Release78Turn an external AI evaluation into a reviewable product decision. The unit of9work is not a headline score; it is the chain from claim to tested system,10harness, budget, validity checks, evidence, limitations, remediation, and11deployment action.1213Independent evaluation can add perspective alongside internal testing, but it14does not become a safety certificate, compliance opinion, adoption signal, or15truth judgment. Keep the evaluator's result, the product's controls, and the16user's outcome as separate evidence layers.1718## When to use1920Use this skill when:2122- a PM is commissioning or reviewing an independent lab, external expert, or23 third-party assessment before deploying an AI model, agent, or product;24- an evaluation needs to support a defined capability, safeguard, or system25 comparison claim rather than an unspecified “quality” claim;26- an agentic result may depend on tools, memory, retries, scaffolding, context27 management, environment state, or a product-facing harness;28- an external report needs a methodology review for scope, independence,29 conflicts, budget, task validity, contamination, refusals, reward hacking,30 evaluation awareness, or broken environments;31- a team must define safe model/checkpoint access, data retention, credentials,32 confidentiality, responsible disclosure, redaction, or publication terms;33- the result may change `Ship`, `Pilot`, `Hold`, `Need evidence`, or `Rollback`34 for a consequential AI capability.3536Use `pm-ai-evaluation-plan` when the primary job is designing a general test37plan before an evaluator or result exists. Use38`pm-ai-review-to-calibration` when the primary job is aligning human or model39judgments. Use `pm-ai-prompt-injection-to-defense` when the primary job is one40untrusted-content attack path and its product controls. Use41`pm-ai-risk-to-control` for a broad hazard register, and42`pm-ai-incident-to-runbook` after a journey-level incident has occurred.4344## Do not use this when4546- the request is to run a live jailbreak, exploit, attack, scanner, or provider47 call;48- the only input is a vendor benchmark and there is no product claim or tested49 system to bound;50- the team wants a generic score, a public ranking, a certification, or a51 promise of production readiness;52- a red-team exercise is being treated as the whole evaluation; red teaming is53 a risk-discovery method and must be scoped separately from capability,54 methodology, and product evaluation;55- the open decision is primarily user approval, authorization, moderation,56 provenance, model migration, cost, or user retention;57- private customer data, hidden prompts, credentials, exploit payloads, or58 chain-of-thought would need to be copied into the public packet.5960## Evidence boundary6162Record each statement as `Observed`, `Reproduced`, `Supplied`, `Proposed`,63`Not run`, `Unknown`, or `Not covered`. A source may describe a method without64proving that the method was applied to this system.6566| Evidence layer | It may establish | It cannot establish by itself |67| --- | --- | --- |68| `provider_or_standard` | documented evaluation, harness, or safeguard capability | this product's result, adoption, or deployment fitness |69| `evaluator_process` | who tested, under what independence, scope, access, and method | that the evaluator had representative coverage or no bias |70| `tested_configuration` | model/version, prompt, tools, memory, environment, budget, and safeguards actually in scope | behavior after a configuration, harness, or policy change |71| `evaluation_result` | observed scores, examples, failures, and analysis under the declared setup | a capability ceiling, real-world prevalence, or generalization beyond scope |72| `product_control` | deterministic gate, approval, sandbox, monitor, or fallback behavior | model capability or user outcome |73| `deployment_outcome` | a verified product run, pilot observation, or user decision at a named surface | broad adoption, safety, or causality without its own evidence |7475If a report omits system identity, harness, budget, denominator, evaluator76conflict, task validity, or limitation, lower the claim or return `Need77evidence`. Never turn a missing field into a passing assumption.7879## Core definitions8081| Term | Working meaning | Do not confuse it with |82| --- | --- | --- |83| `Evaluation claim` | The precise capability, safeguard, or comparison statement the assessment is intended to support | a broad “the model is good” statement |84| `Independent evaluation` | An external assessment whose methods and conclusions are not fully controlled by the product team | a vendor marketing page or internal QA run |85| `Methodology review` | External scrutiny of the framework, data, scoring, or interpretation | a rerun of the product itself |86| `SME probing` | Domain-expert testing of real-world tasks with structured input | representative user research or adoption evidence |87| `Red teaming` | Proactive exploration of risks or attacks to discover issues and build evaluations | a complete deployment evaluation |88| `Harness` | Model-facing prompts, tools, interfaces, memory, retries, validators, control logic, and environment | the model name alone |89| `Maximum elicitation` | Testing for the strongest credible behavior under a defined effort/budget | an unlimited or undisclosed attempt count |90| `Validity hazard` | A condition that can inflate or suppress the measured result, such as reward hacking, refusal, contamination, broken task, or evaluation awareness | ordinary variance |91| `Evidence packet` | Versioned claim, configuration, run, artifacts, limitations, and decision records | a slide with one score |92| `Publication boundary` | What can be shared, redacted, delayed, corrected, or withheld, with an owner | permission to publish every raw artifact |93| `Release decision` | `Ship`, `Pilot`, `Hold`, `Need evidence`, or `Rollback` with conditions and owner | safety, legal, compliance, or adoption certification |9495## Workflow9697### 1. Frame the claim and decision9899Write one sentence:100101> Decide whether evidence from `<evaluation>` supports `<specific claim>` for102> `<user/job or deployment decision>` under `<scope, risk, and date>`.103104State the current workaround, audience, decision owner, false-positive and105false-negative costs, version boundary, and what action the result could change.106Choose one primary claim class:107108- **Capability elicitation:** what the system can plausibly produce under a109 credible setup;110- **Safeguard performance:** how robust a defined control is against a defined111 behavior or adversary model;112- **Controlled comparison:** how systems perform under equivalent conditions;113- **Methodology / assurance:** whether a test design or interpretation is fit114 for the decision;115- **SME probing:** what qualified experts observe on domain tasks.116117Do not combine these into a single score. If the claim is “safe in production,”118split it into model behavior, product controls, user experience, operations,119and deployment evidence.120121### 2. Establish independence, access, and authority122123Create an evaluator ledger before accepting results:124125| Field | Required question | Status |126| --- | --- | --- |127| `evaluator` | Who performed, funded, supervised, and reviewed the work? | `Observed` / `Unknown` |128| `expertise` | Which capability or risk area is in scope, and what qualification supports it? | `Observed` / `Not provided` |129| `conflicts` | What financial, publication, access, or relationship conflicts exist? | `Observed` / `Unknown` |130| `scope` | Which model, checkpoint, product surface, locale, tools, policies, and dates were tested? | `Observed` / `Not provided` |131| `access` | Which credentials, data, retention, logging, chain-of-thought, or reduced-safeguard access was granted? | `Observed` / `Not provided` |132| `authority` | Who can approve test changes, disclose a finding, remediate, or decide release? | `Observed` / `Not provided` |133| `publication` | Which raw artifacts, summaries, limitations, and corrections can be shared? | `Observed` / `Not provided` |134135Treat evaluator access as least privilege. Do not place real customer data,136production secrets, private model weights, or unrestricted external write137authority into a test environment unless the access decision, retention,138isolation, owner, expiry, and rollback are explicit.139140### 3. Freeze the tested system and harness141142Assign stable IDs for the evaluation, model/provider/version, prompt or policy143version, tool and memory versions, task/environment snapshot, harness, budget,144scorer, reviewer, and report. For an agent, record:145146- interface and user-visible surface, not only a raw model endpoint;147- tools, schemas, retries, timeouts, context/compaction, memory, validators,148 approvals, sandbox, and egress boundaries;149- initial state, external data, seed or randomization, locale, and policy;150- turns, tokens, attempts/retries, wall-clock budget, inference cost, and151 expected cost per successful solve where relevant;152- model-level and product-level mitigations that were enabled, disabled, or153 changed during the run.154155If a report used a simpler harness than the real product, state the resulting156claim as a component or lower-bound observation. If the harness changed between157systems or runs, label the comparison `Not comparable` until the difference is158explained or the evaluation is repeated.159160### 4. Define tasks, slices, oracle, and denominator161162Build a small, versioned evaluation matrix before looking at the headline:163164| ID | Claim slice | Task / fixture | Expected behavior | Oracle | Hard blocker | Status |165| --- | --- | --- | --- | --- | --- | --- |166| `E-001` | core user job | safe fictional task | observable completion | deterministic or human reference | none | `Not run` |167| `E-002` | negative route | should abstain or ask | no unsupported assertion/action | route or policy oracle | unsafe action | `Not run` |168| `E-003` | edge / OOD | locale, format, or missing tool | bounded recovery | trace plus review | silent failure | `Not run` |169| `E-004` | validity | scorer or environment trap | no shortcut credit | sample inspection | reward hack | `Not run` |170| `E-005` | safeguard | safe category fixture | control blocks or escalates | product control receipt | privacy/security/action leak | `Not run` |171172Use private or newly constructed tasks when contamination matters. Keep golden,173regression, negative, high-risk, OOD, and holdout slices distinct. Report174counts, missing cases, abstentions, retry/attempt policy, and the denominator.175A rate without its unit, sample, or budget is not a complete result.176177### 5. Check validity hazards before interpreting results178179For every result, record `checked`, `not checked`, or `not applicable` for:180181- **Reward hacking:** the system earns credit through a shortcut in the task,182 scorer, prompt, tool, or harness rather than the intended behavior;183- **Refusals:** safeguards suppress the tested behavior and make capability or184 robustness difficult to interpret;185- **Contamination:** tasks, answers, or close variants were in training data or186 discoverable through tools;187- **Broken problems:** ambiguous ground truth, missing files, flaky services,188 impossible environments, or unfair scoring;189- **Evaluation awareness / sandbagging:** the system changes behavior because190 it recognizes evaluation or the consequences of a high score;191- **Harness and budget drift:** retries, tools, context preservation, scoring,192 or effort differ in ways that change the observed result;193- **Review drift:** a judge, rubric, policy, or human interpretation changed194 after seeing the result.195196Inspect representative trajectories, not only summary scores. A higher score197with more shortcuts, malformed tool structure, or hidden refusal behavior is198not automatically an improvement.199200### 6. Separate result, limitation, and product control201202Write findings in this order:2032041. **Observation:** what happened under the frozen setup.2052. **Interpretation:** what the result may support and what it cannot support.2063. **Limitation:** which task, model, harness, budget, validity, or access gap207 narrows the claim.2084. **Product consequence:** what control, UX state, human review, or scope209 restriction should change.2105. **Re-test:** which artifact or condition must be rerun after remediation.211212Keep a model refusal separate from a product control that prevented a tool call,213blocked an egress path, or gave a user a safe recovery. Keep a capability result214separate from a user outcome. An external evaluator can provide additional215evidence; it does not own the product's final release decision.216217### 7. Set data, confidentiality, and publication boundaries218219For each artifact, define:220221| Artifact | Minimum public form | Restricted material | Owner / expiry |222| --- | --- | --- | --- |223| task set | category, count, sampling rule, safe fixture | private prompts, live credentials, sensitive targets | `Not provided` |224| trace / trajectory | redacted step summary or stable ID | secrets, PII, hidden reasoning, private URLs | `Not provided` |225| score / report | claim, setup, denominator, limitations, result | confidential details, exploit-enabling material | `Not provided` |226| finding | sanitized issue, severity, mitigation, retest state | weaponizable payload or private incident detail | `Not provided` |227228Record data retention, deletion/correction, access logging, responsible229disclosure, accuracy review, publication delay, and correction path. If a result230cannot be independently inspected because of confidentiality, narrow the claim231and label the evidence boundary instead of implying full transparency.232233### 8. Choose the smallest release decision234235- **`Ship`:** the claim is narrow and supported by the correct harness, critical236 validity checks are addressed, no hard blocker is open, ownership and rollback237 are real, and the product controls were tested at the surface that will ship;238- **`Pilot`:** exposure is narrow, high-impact actions are gated, human review or239 safe fallback exists, unresolved limitations have owners and dates, and the240 pilot has an observable stop rule;241- **`Hold`:** evaluator independence, configuration, task validity, denominator,242 access, publication, or control evidence is missing or not comparable;243- **`Need evidence`:** the decision is well framed but the relevant evaluation244 has not run or the supplied result cannot be verified;245- **`Rollback`:** a released claim or control is contradicted by a verified246 critical finding, regression, access failure, or unsafe product behavior.247248For every choice, name the user-visible state, owner, TTL/review date, fallback,249rollback trigger, and one next learning question. Restore requires a new250verification window; silence is not evidence of recovery.251252### 9. Write back and ask one review question253254Return a packet that an evaluator, PM, engineering owner, safety/security255reviewer, and release owner can all inspect. Link the report, configuration,256validity notes, findings, remediation, re-test, publication status, and decision257without copying restricted content. Add a new regression, negative case, or258methodology correction to the authorized evaluation record when appropriate.259260End with one concrete ask, such as:261262- `Approve the narrow pilot with the stated stop rule.`263- `Hold until the harness and denominator are frozen.`264- `Add the validity failure to the holdout and rerun.`265- `Need the evaluator conflict, access, or publication record.`266267## Output contract268269Return these sections in order:270271```markdown272## Evaluation decision on the desk273## Claim and user/job boundary274## Evaluator independence and access ledger275## Tested system and harness276## Task, slice, oracle, and denominator matrix277## Validity hazard review278## Result and limitation ledger279## Product control and remediation280## Data, confidentiality, and publication boundary281## Release, fallback, and rollback282## Not covered283## Review ask284```285286Use `Not provided`, `Unknown`, `Not run`, `Not comparable`, `Not verified`,287`Proposed`, or `Not covered` when the evidence does not support a stronger288statement. For a copy-ready fictional shape, read289`references/independent-evaluation-release-brief.md`.290291## Common rationalizations and red flags292293| Rationalization | Red flag / correction |294| --- | --- |295| “The external lab is credible, so the score speaks for itself.” | Credibility does not replace a claim, configuration, harness, denominator, or validity review. |296| “It is the same model name, so the comparison is fair.” | Tools, retries, context, policy, budget, and product surface can change the result. Freeze or label the comparison. |297| “The benchmark is public and large, so it proves generalization.” | Public or reused tasks can be contaminated; report the tested distribution and limitation. |298| “The model refused, so the product is safe.” | Inspect tool calls, egress, data boundaries, user recovery, and product controls separately. |299| “Red teaming found no issue, so we can ship.” | Red teaming is one risk-discovery slice, not a complete capability, product, or deployment evaluation. |300| “We can publish the raw trace to build trust.” | Redact secrets, PII, hidden instructions, private URLs, and exploit-enabling detail; publish the smallest useful artifact. |301| “The report is confidential, so limitations do not need to be stated.” | Confidentiality narrows the claim; it does not justify silent uncertainty. |302| “A green CI run makes the external assessment complete.” | CI proves repository checks, not evaluator independence, harness validity, product behavior, or adoption. |303304## Edge cases305306- **No external evaluator:** produce a commissioning brief with `Result: Not307 run` and keep the decision at `Need evidence`.308- **Vendor-produced benchmark:** record it as `Supplied`, name the vendor's309 interest and setup, and do not label it independent.310- **Model-only test for an agent product:** state the product harness gap and311 do not generalize the model result to the agent.312- **Different harnesses:** report separate results or rerun under a shared313 setup; do not rank systems from an unqualified comparison.314- **Reduced safeguards or early checkpoint:** record the exact access boundary,315 isolation, retention, and the fact that the result may not represent the316 deployed configuration.317- **High score with reward hacking or broken tasks:** quarantine the result,318 correct the task/scorer, and return `Hold` or `Need evidence`.319- **High-risk red-team finding:** keep the safe summary public, restrict the320 payload, assign remediation, and do not claim the absence of a finding means321 the system is secure.322- **Small sample or no denominator:** keep the result descriptive and label the323 decision `Not measurable` or `Need evidence`.324- **Confidential third-party report:** preserve the report locator, scope,325 methodology review, and limitation statement without exposing restricted text.326- **Fictional or synthetic fixture:** label it clearly; it supports workflow327 shape and known edge cases, not real capability, safety, user, or adoption.328329## Final check330331Before returning the packet, confirm:332333- the decision, user/job, claim class, evaluator, system, version, harness,334 budget, task set, denominator, and observation window are visible or marked335 missing;336- independence, conflict, access, data retention, confidentiality,337 publication, and responsible-disclosure boundaries are explicit;338- capability, safeguard, comparison, methodology, product control, and user339 outcome evidence are not collapsed into one score;340- golden, negative, edge/OOD, validity, high-risk, and holdout slices are341 represented where relevant;342- reward hacking, refusals, contamination, broken problems, evaluation343 awareness, harness/budget drift, and review drift are checked or marked;344- product controls, human review, fallback, rollback, owner, TTL, and stop345 condition are observable;346- the report distinguishes supported, unsupported, unknown, not run, blocked,347 and not covered claims;348- no secret, private payload, raw customer content, hidden reasoning, or349 unsupported safety, truth, legal, adoption, or production claim was added.350351## References and first run352353- For a copy-ready fictional workflow, read354 [`examples/first-run.md`](examples/first-run.md). It is a **fictional355 fixture**, not a live evaluation or deployment result.356- For a complete packet shape and source ledger, read357 [`references/independent-evaluation-release-brief.md`](references/independent-evaluation-release-brief.md).358- For current technical source context, read the official sources in the359 reference file; refresh them before making version-sensitive claims.360361## Not covered362363- This skill does not run an evaluation, red team, jailbreak, exploit, scanner,364 model call, external provider, or third-party engagement.365- It does not certify safety, security, compliance, legal status, truth,366 authorship, production readiness, user satisfaction, adoption, or GitHub367 growth.368- It does not authorize access to private data, model weights, credentials,369 hidden reasoning, production tools, or external write surfaces.