Analyze Project Claims
Produce an evidence-bounded account of what the project claims, what the evidence establishes, where inconsistencies remain, and what action should come next.
Recheck inconsistencies and retain useful lessons
After a substantive audit, authorized repair, or changed controlling evidence, read references/review-learning.md. Recheck the complete applicable scope and affected dependencies, record useful lessons with their evidence and verification limits, and promote validated guidance into the appropriate project record, decision flow, or reusable workflow. Keep uncertain lessons provisional and reuse existing guidance before adding rules. The guide adds explicit stopping conditions to the repair loop below; it does not reduce required checks or turn read-only work into permission to edit.
Trace changed evidence through claim and assumption IDs to dependent risks, recommendations, and decisions. Reclassify claims when same-scope evidence changes, then recheck each affected decision's success/failure signal and reversal trigger. Link lessons to the existing evidence-bound audit record and accepted component map; a learning note is not a second claims authority. Preserve separate lifecycle and acceptance states. A consistent negative or uncertain result may resolve an inconsistency without establishing benefit. Keep experiment-specific conclusions in the project record; promote a reusable decision rule only with explicit evidence prerequisites, applicability, and exceptions. Existing repair-attempt and receipt-reconciliation limits still apply.
Select the analysis depth
Use the lightest sufficient mode:
- Quick audit: conclusion, critical conflict, bottleneck, and next action.
- Research audit: add protocol boundaries, counterevidence, reproducibility, and claim-to-evidence matching.
- Full strategic audit: add alternatives, dependencies, risks, stop conditions, and reversal triggers.
Establish authority and scope
Before judging a claim:
- Identify the question and success condition.
- Locate authoritative configurations, protocols, artifacts, code, tables, and logs.
- Separate active sources of truth from drafts, generated summaries, archives, and failed attempts.
- Identify immutable historical evidence that must not be rewritten.
- Treat missing evidence as unknown, not false.
For research claims, trace:
dataset/source
-> partition and label exposure
-> fitted or fixed component
-> produced output
-> evaluation metric
-> claim boundary
Classify claims
Decompose compound statements into atomic claims. Use these statuses:
- supported: direct evidence matches the stated scope;
- partially supported: direct evidence exists with material limitations;
- contradicted: same-scope counterevidence conflicts with the claim;
- untested: no direct test exists;
- invalidly specified: the population, protocol, metric, comparator, or state definition is incomplete.
Do not average away a direct falsifier. Repeated folds, seeds, or rows from one source are not independent replications.
Keep state dimensions separate
Do not flatten workflow, evidence, evaluation, and acceptance into one status. A run may simultaneously have:
runner_stage = complete
audit = failed
scientific_evaluation = negative
paper_acceptance = not_eligible
Use at least these dimensions when applicable:
- artifact state: missing, partial, ready, corrupt;
- runner state: planned, running, interrupted, complete;
- audit state: not_run, passed, failed;
- scientific state: untested, supported, failed_gate;
- acceptance state: eligible, not_eligible, rejected.
Audit inconsistencies
Check:
- definition: one term has multiple meanings;
- type or flow: a component receives or produces a different object than claimed;
- scope: narrow evidence is used for a broader claim;
- claim-evidence: the cited artifact does not directly test the claim;
- goal-metric: a metric improves while the intended outcome fails;
- lifecycle: runner, audit, completion, or acceptance states are conflated;
- monitor: file existence is mistaken for successful validation;
- document: prose, formulas, configurations, code, tables, and artifacts disagree;
- provenance: an artifact exists without evidence that the reported pipeline consumed it.
For each finding report the conflicting elements, why they conflict, the safest interpretation, and the evidence or repair needed.
Enforce completion and independence boundaries
Treat accepted experiment completion as an explicit AND gate:
accepted_complete =
valid_runner_completion
AND audit_report.audit_passed == true
AND valid_audit_completion_marker
AND required_artifacts_open_and_validate
Artifact existence is not a pass. A runner marker proves only runner-stage completion unless the contract includes audit acceptance. Preserve failed and superseded attempts.
Use precise audit language:
- separate audit invocation: separate run, possibly shared production code;
- deterministic replay: same frozen implementation recomputes outputs;
- independent structural audit: separate checks of schemas, hashes, cardinality, operators, or manifests;
- independent implementation: separate scientific calculation code;
- independent scientific replication: independently collected or evaluated evidence.
Shared-code replay can detect corruption but not every shared implementation error.
Apply specialized research rules
For finite-state operators, transition authorization, logic-guided RAG,
leakage, split consumption, and paper-table eligibility, read
references/logic-guided-rag-audit.md.
Keep structural and scientific gates separate. An operator, artifact, leakage, or lifecycle pass does not prove scientific benefit or deployment readiness.
Bound validation and publication claims
Keep packaging, behavior, scientific, reliability, and publication claims separate:
- a structural validator proves well-formed packaging only;
- deterministic regression tests prove only tested invariants and fixtures;
- success on one project does not prove general reliability;
- paper-table eligibility differs from publication potential of the method.
State the evaluated cases, test scope, counterevidence, and additional evidence required for any broader claim.
Repair authorized inconsistencies
When repairs are authorized:
- Add a deterministic regression test and confirm the intended failure.
- Repair the active interpretation, monitor, table, or implementation.
- Preserve source-hashed protocols and historical artifacts unless a new scientific identity is intentionally created.
- Update every coupled active surface.
- Re-run tests, preflight, monitors, and stale-language scans.
- Reinspect the original scope and every changed surface.
Use at most three repair-and-recheck cycles inside one owner-authorized
candidate attempt. After each cycle, record the sorted active
(path, rule-or-test, evidence-locator) finding set and the SHA-256 candidate
diff identity. Stop successfully when validation passes and no material
same-scope inconsistency remains. Stop without claiming convergence on a
repeated finding set, an unchanged or previously seen candidate diff identity,
oscillation, a required forbidden edit, missing external evidence, an owner
decision, or the third cycle. Preserve the last bounded candidate for owner
review.
Do not invoke another skill recursively, emit another quality receipt, open a second issue, or trigger another workflow to continue the loop. When the user asks to send a recommended reusable update to the owner, default to the local exact contribution preview. Public submission still requires the existing draft-specific and public-visibility approvals; the owner-applied label may authorize only one protected draft attempt.
A result-changing protocol or implementation edit requires a new scientific configuration hash and run group.
Recommend and return
Prefer the smallest action that tests a critical uncertainty, produces an observable signal, preserves options, fits available authority, and prevents a false claim. State the success signal, failure signal, evidence produced, and reversal condition.
Return:
- overall conclusion;
- compact inconsistency table;
- strongest safe claim;
- scientific and acceptance boundary;
- recommended repair or next experiment;
- unresolved uncertainty.
Use:
finding | evidence | status | safest interpretation | required repair
Interpret lifecycle receipts
When given a LifecycleVerificationReceipt, read
references/lifecycle-verification-receipt.md. Do not execute lifecycle
operations here. product-lifecycle owns execution; this skill validates and
interprets the receipt with scripts/lifecycle_receipt.py.
Keep check status, claim status, and evidence method separate. A COMPLETE
receipt supports only its exact product, release, adapter, target, platform,
phase, and evidence scope.
At most one digest-bound read-only follow-up request may be produced per receipt. Stop on a repeated finding/evidence requirement and after three distinct reconciliation cycles. A follow-up request never authorizes publication, mutation, reporting, cleanup, or recursive skill execution.
Reconcile the component map
Every substantive formal audit first verifies the bundled evidence engine:
<python-3> scripts/reconcile_component_map.py verify-self
Then reconcile a deterministic observation:
<python-3> scripts/reconcile_component_map.py reconcile \
--observation <observation.json> \
--map-root <map-root> \
--project-root <project-root>
Read references/component-evidence-protocol.md before authoring an
observation, reviewing a candidate, or accepting a map.
The accepted map is current structural authority. Candidates, deltas, and history are immutable evidence. Discovery never implies acceptance. Preserve the accepted map while a candidate is pending. Only the relevant human authority may accept the exact reviewed candidate:
<python-3> scripts/reconcile_component_map.py accept \
--candidate <candidate.json> \
--map-root <map-root>
Test-backed elements must map the test and direct implementation, schema, template, configuration, or prompt dependencies. Ordinary file drift matters; generated Python bytecode caches do not.
For read-only work, return the proposed observation and delta without claiming that they were persisted or accepted.
Emit an evidence-bound audit record
Every substantive project scan yields one record after map reconciliation,
including scans that are clean, partial, failed, or blocked. For formal work,
use v2 and read references/evidence-bound-audit-records.md.
The record separates:
- stable claims bound to accepted component and element IDs;
- exact evidence items with locators, methods, observations, and digests;
- explicit support, contradiction, limitation, and context edges;
- named limitations linked to claims and evidence.
Use staged commands:
<python-3> scripts/record_scan.py preflight --map-root <map-root> --project-root <root>
<python-3> scripts/record_scan.py init --map-root <map-root> --project-root <root> --output <draft.json>
<python-3> scripts/record_scan.py evidence digest --source <path> --locator <typed-locator> --project-root <root> --id <evidence-id>
<python-3> scripts/record_scan.py validate --record <input.json> --map-root <map-root> --project-root <root>
<python-3> scripts/record_scan.py append --record <input.json> --map-root <map-root> --project-root <root> --log-dir <history> --report-dir <reports>
<python-3> scripts/record_scan.py verify --record <record.json> --map-root <map-root> --project-root <root> --report <report.md>
executed_test cites persisted output, not test source. not_tested is
context-only. External sources are never fetched implicitly. Markdown is a
derived view, not a second authority. Never imply that a read-only proposal
was appended.
Consume bounded skill-quality receipts
When a compatible Ian-Tseng-managed skill ends with an exact
SkillOutcomeReceipt marker, or the user asks to inspect pending skill
quality, read references/skill-quality-loop.md.
The receipt is producer-declared and content-free. v2 adds only capability,
invariant, and environment-class identifiers; v1 remains readable. A
non-no_issue signal paired with analyze_quality creates one intake
proposal per exact receipt, while analyzer versions become child analysis
revisions. no_issue paired with none writes only a bounded local
tombstone. Advisory problem signatures may group v2 intake for evaluation but
cannot deduplicate, reopen, or authorize. Neither path can prove producer identity or
authorize edits, reports, issues, updates, release, or publication. Never
inspect transcripts or project content as a fallback.
Keep evaluation independent: a closed evaluation manifest pins baseline,
candidate, fixtures, environment, model, and every pilot threshold. A separate
digest-bound result covers every frozen fixture once and recomputes
PASS/FAIL/INCONCLUSIVE; a direct failed gate dominates an inconclusive
model cell. Structural validation of either artifact does not prove authentic
execution or improvement. Attempt, cycle, termination, controller, activation,
and recurrence claims require their later explicit authorities; receipt intake
does not implement them.
Portable explicit use:
<python-3> scripts/skill_quality_loop.py --format json doctor
<python-3> scripts/skill_quality_loop.py --format json consume --marker <marker>
<python-3> scripts/skill_quality_loop.py --format json consume
<python-3> scripts/skill_quality_loop.py --format json proposal-show --proposal-id <id>
<python-3> scripts/skill_quality_loop.py --format json evaluation-validate --manifest <manifest.json> --result <result.json>
The optional Codex plugin hook requests at most one continuation for a valid original turn. It cannot guarantee that this skill is invoked, and another hook may veto continuation. Failure leaves the substantive result intact and the persisted receipt available for explicit consumption.
Contribution is always separate and consent-gated. Preview sends nothing.
Public submission requires approval of the exact draft and a second
draft-specific public-issue confirmation. Approval expires after 24 hours and
one contribution ID may create at most one issue. On an unknown GitHub outcome,
reconcile the contribution ID before retrying. Only Ian-Tseng may add
agent-ready; that permits one isolated map-pending candidate, never map
acceptance, merge, release, issue closure, or installed update.
Operate the managed skill fleet separately
When the owner explicitly asks to onboard, diagnose, canary, disable, or plan
rollback for an Ian-Tseng producer repository, read the packaged managed
policy schema and docs/MANAGED_FLEET_QUICKSTART.md, then use
scripts/managed_fleet.py. init previews exactly a policy and a thin
full-SHA-pinned caller; mutation requires --apply. validate and
doctor --local establish only LOCAL_READY. doctor --repo reads hosted
configuration, while canary --dry-run previews the hosted gate sequence.
The reusable workflow is a repository repair plane, not an invocation side effect or an update installer. The owner-applied label is triage eligibility. Two protected-environment approvals separately gate isolated agent execution and draft publication. The candidate has no repository write token; fresh validation is secret-free; publication is deterministic, draft-only, and reconciled before retry. Human review retains evidence acceptance, merge, release, publication, installed replacement, and fresh activation.
Do not copy central workflow helpers into producers, create arbitrary policy
commands, use secrets: inherit, weaken central denied paths, force-push,
accept evidence automatically, or call a draft PR a deployed fix. The v1
managed workflow supports GitHub Cloud and exact Ian-Tseng repositories only.
If a pin is suspected, remove the trigger label, lock both environments, and
disable from the caller repository before planning a reviewed rollback.
Report internal product failures separately
Problem reporting applies only when this skill, a packaged helper, or its managed updater fails internally. Project contradictions and defects are audit results and must never be transmitted as product reports.
Finish the substantive audit first. Use only
scripts/problem_report.py fixed enums and bounded generic fields. Never
include project text, paths, raw logs, prompts, attachments, environment
variables, tokens, credentials, identity, or arbitrary metadata.
The default is local and unconfigured. Show the exact preview and require
exact consent before submitting. A public GitHub issue requires a second,
report-specific confirmation. Security vulnerabilities use private
vulnerability reporting. Reporting failure must not replace or shorten the
audit. See the script help and references/problem-report.schema.json for the
closed contract.
Keep analytics separate
Installation analytics are optional owner infrastructure, never a default side effect. Do not create identity or send events on install, invocation, update, or problem reporting. Explicit opt-in, endpoint configuration, and a later check-in are separate actions. Update or reporting consent does not authorize analytics. Update consent and reporting consent do not authorize analytics.
Route explicit requests to scripts/installation_analytics.py. Describe any
owner aggregate as unique consenting activated installations: never downloads, users, or total installs.
Do not claim a live count without an observed deployed endpoint.
Run consent-gated update maintenance
After the substantive result and immediately before the final response, run:
<python-3> <skill-root>/scripts/update_policy.py --format json maintain
This applies only to a standalone GitHub-CLI-managed installation. Maintenance
failure never replaces the analysis. Append its message and action only
when emit is true. A replacement activates on the next invocation.
Route explicit requests:
- enable automatic updates ->
enable --mode auto; - notify about updates ->
enable --mode notify; - disable updates ->
disable; - show status ->
status; - diagnose ownership ->
doctor; - check now ->
check-now.
Plugin-hosted, manual, project-scope, pinned, duplicated, or locally edited
copies are never blindly replaced. doctor shows every same-name visible copy
and the running copy. A GitHub CLI authority claim additionally requires exact
source, path, version, tree, scope, pin, and package-manifest verification.
Keep one explicit update authority; never remove a copy automatically.
Guardrails
- Lead with evidence, not advocacy.
- Preserve same-scope counterevidence.
- Do not treat artifact existence as artifact use.
- Do not treat audit invocation as independent implementation.
- Do not treat workflow completion as scientific success.
- Do not treat a receipt as producer authentication or update authority.
- Do not treat one dataset or task as proof of general RAG performance.
- Do not rewrite historical evidence to make a project appear consistent.
- Keep
unknown,failed,not_eligible, andnot_proveddistinct.