/find-implicit-state
C# semantic branch
Run the sibling _csharp-semantic provider from its guide, then enter through
scripts/detect_csharp_state.py; knowledge/csharp-v1.md gives the exact
consumer command. A candidate is compiler-resolved selected-source evidence,
not proof of a closed domain. Serialization, reflection, generated inputs,
external callers, frameworks, and binary compatibility remain unresolved.
Kotlin/JVM 2.4.10 branch
Trigger this branch only for an exact kotlin-semantic-project.json target.
Keep sibling _kotlin-semantic, read
../_kotlin-semantic/GUIDE.md, produce its
pinned fact pack, then enter through scripts/detect_kotlin_state.py. Direct
literal writes to one resolved String state/status/phase property are a
review lead, not a closed-domain proof; the deprecated K1 API is not a stable
Analysis API. Delegates/custom setters, overrides,
reflection, generated/plugin sources, Gradle variants, Java, frameworks,
runtime writes, enum safety, and mutation remain unresolved.
Swift 6.3.3 semantic branch
Use scripts/detect_swift_state.py only with one current complete
swift-semantic-facts-v2 pack from sibling _swift-semantic-readonly. It
reports compiler-resolved direct literal writes and comparisons for one exact
selected-target String state/status/phase property, then requires a
candidate-hash-bound human verdict. Observed literals do not prove a closed
runtime domain, serialization compatibility, framework/dynamic consumers,
generated or conditional variants, enum safety, or mutation authority.
C++20 branch
Use scripts/detect_cpp_state.py with _cpp-semantic; run the script with
--help for the exact CLI. It reports exact resolved string writes to one
namespace-qualified field under a current complete C++20 compiler-owned graph.
Overloads remain distinct; closed-domain, alias/callback, ODR/ABI,
specialization, dynamic-dispatch, external-variant, and migration claims are
refused.
C17 branch
Use scripts/detect_c_state.py with the sibling _c-semantic provider; run
python3 scripts/detect_c_state.py --help for the exact CLI. This external-
library branch emits one human-reviewed enum_review_only candidate. Aliasing,
callbacks, external mutation, macros, variants, closed-domain proof, and
automatic migration remain unresolved.
PHP and Ruby
For a selected PHP or Ruby run, load ../_php-semantic/GUIDE.md or
../_ruby-semantic/GUIDE.md. Both branches require candidate-hash-bound human
review and never claim a closed runtime domain or safe enum conversion.
Dart v1
Dart v1 uses the sibling map-subsystem SDK-LSP provider to nominate a
direct class String state, status, or phase field only when at least
three literal assignments or comparisons resolve to that exact declaration.
It requires a candidate-hash-bound human review before a finding is promoted;
the literals never prove that the value domain is closed. Copy sibling
map-subsystem with this skill.
SKILL_ROOT=".agents/skills/on-demand/find-implicit-state"
python3 "${SKILL_ROOT}/scripts/detect_dart_state.py" \
--project-root "$PWD" --target lib \
--output-dir "$PWD/reports/implicit-state/dart"
Typed state, insufficient evidence, local homonyms, dynamic access,
serialization/wire carriers, tests, generated code, examples, vendor code,
reflection, isolates, external compatibility, and Flutter state are excluded
or unresolved. The first run normally stops at human_review_required; pass
an accepted review directory with --reviews-dir to produce final findings.
Rust v1
Rust v1 proposes repeated bare string operations on a bounded state field for
human /extract-enum review. It never claims the domain is closed and refuses
promotion through typed enums, insufficient evidence, traits/generics,
macros, cfg, unsafe/FFI, or excluded roles. Copy sibling map-subsystem.
SKILL_ROOT=".agents/skills/on-demand/find-implicit-state"
python3 "${SKILL_ROOT}/scripts/detect_rust_state.py" \
--project-root "$PWD" --target . \
--output-dir "$PWD/reports/implicit-state/rust"
You are the orchestrator for an implicit-state audit. Your job is to drive a pipeline of detectors and sub-agent verifiers; the bucket judgment calls live in the scout brief and the knowledge files, not in this prompt.
The two sub-patterns (stringly-typed state, tuple-inferred identity),
the four buckets (extract_enum_candidate, introduce_fk_candidate,
enum_already_used, legacy_allow_list), and the verification checklist
are documented in knowledge/verification.md — scouts read it, you
don't.
How success is judged
- Every reported candidate carries a Stage 3 scout verdict at
scout/<candidate_id>.jsonin one of the four buckets (extract_enum_candidate/introduce_fk_candidate/enum_already_used/legacy_allow_list), with the hit evidence attached — nothing reachesreport.mdungraded. - The closeout pastes the real Stage 1/2/4 stderr lines
(
[detect_implicit_state],[collapse_implicit_state],[report_implicit_state]) plus the scout JSON count; claims without those artifacts do not satisfy the audit. - Each actionable candidate is routed to its named handoff:
/extract-enum <symbol>or/introduce-fk <symbol>. - Zero edits to production code — detection-only audit. Write toward these gates from Stage 0.
TypeScript closed-state branch
Use this branch only for .ts / .tsx code whose first-party state/status/phase receiver resolves to a closed string-literal union or project-native enum. Its supported outcome is evidence for replacing repeated first-party bare literals with an exported as const runtime value object and a derived union type. It does not claim a TypeScript ORM, migration, tuple-identity, or general text-literal detector.
The bounded syntax contract covers direct property comparisons, reversed
comparisons, one-hop local const aliases initialized directly from the
property, plain and ??= assignments, and every property target in a chained
assignment. Vendor attribution comes from an explicit semantic receiver type
named Vendor*Payload|Request|Response|Event|Message|Wire; filenames and
nearby text never establish a vendor boundary. Parentheses are transparent
around state operands, direct alias initializers, literals, and chained
assignment expressions. Computed properties, other assertion wrappers, and
general dataflow remain out of scope. Invalid TypeScript exits 2.
Host prerequisites: the target project owns a compatible typescript package and a readable tsconfig.json. The launcher resolves typescript from that project's package.json; it never uses a toolkit, global, or downloaded compiler. A missing package or tsconfig is a clear exit 2, not a lexical fallback.
Run this branch instead of Python Stages 1–4:
REPORT_DIR="reports/implicit-state/scan-typescript-$(date +%Y%m%d-%H%M%S)"
mkdir -p "$REPORT_DIR"
node .claude/skills/find-implicit-state/scripts/detect_typescript_state.mjs \
--target "$(pwd)" \
--project-root "$(pwd)" \
--tsconfig "$(pwd)/tsconfig.json" \
--output "$REPORT_DIR/findings.jsonl"
Grade the run from the emitted JSONL and stderr line [detect_typescript_state]: it must contain first-party operations only for the migration candidate, while retaining classification records for typed authorities, vendor wire boundaries, tests/fixtures, unrelated status text, and open-ended strings. Hand that exact JSONL to the TypeScript branch of /extract-enum; do not run the Django collapse/scout/report stages on it.
Checked JavaScript closed-state branch
Use this branch only with a host-local typescript Compiler API and an
explicit jsconfig.json or tsconfig.json that enables both allowJs and
checkJs. It accepts .js, .jsx, .mjs, and .cjs, but promotes an
operation only when the receiver has a demonstrated finite JSDoc authority;
untyped/open strings remain classification evidence, never a migration lead.
The manifest records the config, diagnostics, unresolved modules, uncovered
sources, compiler-parsed JSDoc, and TypeChecker inference. A missing tool or
config is unsupported, malformed selected JS is syntax-error, and any
unresolved/excluded source is partial. Do not use npx, a global compiler,
or framework inference.
: "${TARGET:?Set TARGET to the checked-JavaScript directory to inspect}"
JSCONFIG="${JSCONFIG:-jsconfig.json}"
SKILL_ROOT=""
for SKILL_CANDIDATE in \
".agents/skills/on-demand/find-implicit-state" \
".agents/skills/find-implicit-state" \
".claude/skills/find-implicit-state"
do
if [ -f "${SKILL_CANDIDATE}/SKILL.md" ]; then
SKILL_ROOT="$(cd "${SKILL_CANDIDATE}" && pwd)"
break
fi
done
if [ -z "${SKILL_ROOT}" ]; then
printf '%s\n' "find-implicit-state is not installed in .agents/skills/on-demand, .agents/skills, or .claude/skills" >&2
exit 2
fi
node "${SKILL_ROOT}/scripts/detect_typescript_state.mjs" \
--target "${TARGET}" --project-root "$(pwd)" --tsconfig "${JSCONFIG}" \
--output reports/implicit-state/javascript.jsonl \
--manifest reports/implicit-state/javascript.manifest.json --language javascript
This is a standalone host-root command: it resolves the selected skill itself under either supported installation layout.
Go implicit-state review branch
For a Go target, read and follow knowledge/go-v1.md. Load that file only for
Go work. Its result is a review candidate—not compiler proof that the domain is
closed—and only its first_party_state_operation records may proceed to
/extract-enum.
Java 17 direct-String review branch
Use this branch only for a JDK 17+ host with java and javac on PATH. It
uses the bundled compiler-tree/type helper, not Maven, Gradle, JARs, a network
download, or a shared Java platform. It promotes only a compiler-resolved,
top-level direct java.lang.String field named state, status, or phase
with at least three direct assignments / String.equals / Objects.equals
operations and two distinct literals. That is review evidence, not proof that
the domain is finite. Aliases, getters/setters, switches, dataflow, ORM
converters, serialization behavior, nested owners, and Kotlin are out of
scope.
field == "value" or != on that String field is an
unsafe_string_comparison correctness finding. It never counts toward an enum
candidate. Generated, test, vendor-path, semantic vendor-wire, unrelated, and
low-evidence shapes remain explicit classifications.
: "${TARGET:?Set TARGET to the Java directory to inspect}"
SKILL_ROOT=""
for SKILL_CANDIDATE in ".agents/skills/on-demand/find-implicit-state" ".agents/skills/find-implicit-state" ".claude/skills/find-implicit-state"; do
if [ -f "${SKILL_CANDIDATE}/SKILL.md" ]; then SKILL_ROOT="$(cd "${SKILL_CANDIDATE}" && pwd)"; break; fi
done
if [ -z "${SKILL_ROOT}" ]; then printf '%s\n' "find-implicit-state is not installed" >&2; exit 2; fi
REPORT_DIR="reports/implicit-state/java-state-$(date +%Y%m%d-%H%M%S)"
python3 "${SKILL_ROOT}/scripts/detect_java_state.py" --target "${TARGET}" --project-root "$(pwd)" --output "${REPORT_DIR}/hits.jsonl" --findings "${REPORT_DIR}/findings.json" --report "${REPORT_DIR}/report.md" --scan-id "$(basename "${REPORT_DIR}")" || exit $?
ln -sfn "$(basename "${REPORT_DIR}")" reports/implicit-state/latest
Hand exactly one status: accepted, bucket: extract_enum_candidate finding
from the complete findings.json to /extract-enum; do not give the next
skill the source tree to re-detect. The final report.md and findings.json
are the review artifacts. findings.json fingerprints every selected Java
source so the proposal consumer can reject stale field or caller evidence.
Syntax or missing/old-JDK failures exit 2 without publishing them; unresolved
or symlink evidence is partial, never clean.
Scope
- Target path: the required
--targetargument. Must be a directory. - Project root: this worktree's root.
- Python:
.venv/bin/pythonin this repository; the bundled detector, collapse, and reporter are stdlib-only and run with hostpython3after a stock installation. - Project-specific defaults (known enums, tuple-identity hot
spots, noqa conventions, detection gaps): in
knowledge/.
Pipeline stages (each has a contract)
Each stage reads files the previous stage wrote and writes files the
next stage reads. Run scripts with .venv/bin/python and capture stderr so
failures surface.
Stage 0 — Setup
Pre: none. Post: ${REPORT_DIR} exists, latest symlink
points to it.
TS=$(date +%Y%m%d-%H%M%S)
REPORT_DIR="reports/implicit-state/scan-${TS}"
mkdir -p "${REPORT_DIR}/scout"
ln -sfn "scan-${TS}" reports/implicit-state/latest
Stage 1 — Detect
Pre: target directory exists. Post: ${REPORT_DIR}/hits.jsonl
present.
.venv/bin/python .claude/skills/find-implicit-state/scripts/detect.py \
--target <target> \
--project-root "$(pwd)" \
--output "${REPORT_DIR}/hits.jsonl"
The detector surfaces four hit shapes: stringly_compare,
stringly_field, possible_state_literal, tuple_identity. It does
not resolve cross-module enum references; the scout does that at
Stage 3.
Stage 2 — Collapse
Pre: hits.jsonl. Post: ${REPORT_DIR}/candidates.jsonl —
one record per (file, pattern) bucket with hit count, confidence
tier, fields/symbols touched, and a recommendation hint.
.venv/bin/python .claude/skills/find-implicit-state/scripts/collapse.py \
--hits "${REPORT_DIR}/hits.jsonl" \
--output "${REPORT_DIR}/candidates.jsonl"
Confidence tiers:
high— everytuple_identityhit, or anystringly_comparegroup with ≥3 hits on the same field in one file.medium—stringly_fieldgroups,stringly_comparebelow the ≥3-same-field threshold,possible_state_literalgroups in files that also have astringly_comparehit.low—possible_state_literalin files with no compare hits.
Stage 3 — Verify (parallel fan-out)
Pre: candidates.jsonl. Post:
${REPORT_DIR}/scout/<candidate_id>.json for every verified candidate.
This is the only stage where LLM judgment runs. You do not verify candidates yourself — dispatch one sub-agent per candidate. Each sub-agent receives:
- the candidate JSON (one line from
candidates.jsonl), - the prompt template from
agents/verify.md, - paths to the
knowledge/*files, - an output path it must write to.
Budget: verify up to 10 high-confidence candidates by default, prioritizing in this order:
tuple_identityhits first — rare and load-bearing.stringly_comparegroups with ≥3 hits on the same field.stringly_fieldgroups (model declarations).- Remaining medium-confidence groups.
If the user asked for a deeper scan, raise the budget. If the user asked for a specific file, filter candidates to that path before dispatch.
For each selected candidate, expand agents/verify.md (substitute
{{candidate_id}}, {{candidate_json}}, {{project_root}},
{{skill_root}}, {{output_path}}) and dispatch with
subagent_type=general-purpose. Send all Agent calls in a single
message so they run concurrently.
Declare the verdict to every scout: its output is accepted only if it
writes valid JSON at {{output_path}}, uses one of the four buckets,
preserves the candidate_id, and includes the hit evidence fields from
the schema. When merging, reject or re-dispatch malformed scout files;
do not let report.py turn an unverified candidate into an actionable
handoff.
If a scout returns invalid JSON, re-dispatch once with a stricter "respond only with file-write confirmation" nudge; skip the candidate if it fails twice.
Dispatch mode — Agent tool vs cheap subprocess
This skill declares scout_model: cheap — the verify step is read-and-
classify against the four buckets in verify.md (extract_enum_candidate
/ introduce_fk_candidate / enum_already_used / legacy_allow_list).
The scout reads the enclosing function and one noqa-grep, no cross-file
synthesis, no shell. Safe on Haiku-class scouts.
For nesting-safe + low-cost fan-out, dispatch each candidate as a
tools/code_agent.py --read-only subprocess via
.claude/skills/_common/dispatch_scout_cheap.sh. The --read-only
flag drops bash, spawn_agent, claude_tools, and validate_jsonld — the
scout has only read_file/write_file/glob/grep, with workdir
containment enforced (commit 168ca3c1). Cheap models can't
hallucinate calls to tools that aren't in the registry.
# One subprocess per candidate; parallelize with `&` + wait.
while read -r line; do
cid=$(jq -r '.candidate_id' <<<"$line")
out="${REPORT_DIR}/scout/${cid}.json"
.claude/skills/_common/dispatch_scout_cheap.sh \
.claude/skills/find-implicit-state/agents/verify.md \
"$out" \
candidate_id="$cid" \
candidate_json="$(jq -c . <<<"$line")" \
project_root="$(pwd)" \
skill_root=".claude/skills/find-implicit-state" \
output_path="$out" &
done < "${REPORT_DIR}/candidates.jsonl"
wait
Tradeoffs. Cheap subprocess dispatch adds 2-4s spawn per scout
and runs Claude Haiku 4.5 through the team's Expedient gateway by
default (see 0s spawn) and
uses the orchestrator's session model (Sonnet/Opus tier — more
judgment, billed). Use the cheap subprocess by default; fall back to
tools/agent-config.json); set DISPATCH_SCOUT_MODEL
to swap in any other alias (e.g., cerebras for personal-account
free-tier capacity). The Agent tool path is faster (Agent when (a) only a handful of candidates need verification
interactively and the user is watching, or (b) a tuple-identity
candidate is genuinely ambiguous between identity and freshness usage
and warrants the better model.
Stage 4 — Report
Pre: candidates.jsonl, scout/*.json. Post:
${REPORT_DIR}/report.md and ${REPORT_DIR}/findings.json.
.venv/bin/python .claude/skills/find-implicit-state/scripts/report.py \
--scout-dir "${REPORT_DIR}/scout" \
--candidates "${REPORT_DIR}/candidates.jsonl" \
--output-md "${REPORT_DIR}/report.md" \
--output-json "${REPORT_DIR}/findings.json" \
--scan-id "scan-${TS}" \
--target <target>
# Effectiveness log — one line per run. Buckets come from findings.json.
.venv/bin/python scripts/log_effectiveness.py \
--skill find-implicit-state \
--scan-id "scan-${TS}" \
--target <target> \
--findings-total "$(.venv/bin/python -c 'import json,sys; d=json.load(open(sys.argv[1])); print(d.get("summary",{}).get("findings_total", len(d.get("findings", []))))' "${REPORT_DIR}/findings.json")" \
--buckets "$(.venv/bin/python -c 'import json,sys; print(json.dumps(json.load(open(sys.argv[1])).get("summary",{}).get("buckets", {})))' "${REPORT_DIR}/findings.json")"
Stage 5 — Summarize
Report to the user in ≤10 lines:
- counts by bucket (extract_enum_candidate / introduce_fk_candidate / enum_already_used / legacy_allow_list),
- counts by pattern (stringly_compare / stringly_field / possible_state_literal / tuple_identity),
- top 3 candidates (one line each: file, pattern, recommendation),
- path to
${REPORT_DIR}/report.mdand thelatestsymlink, - recommended next slash command (
/extract-enum <symbol>,/introduce-fk <symbol>, or/find-implicit-stateagain after cleanup).
The report is the source of truth — do not enumerate every candidate.
Replay case
When detect.py, collapse.py, report.py, or the scout JSON schema
changes, replay a disposable Python target with one model containing
status = models.CharField(...) without TextChoices, one function that
compares job.status == "pending", and one tuple-identity
.filter(status=..., *_at__...).first() shape. Expected evidence:
Stage 1 writes hits.jsonl with both stringly-state and tuple-identity
hits; Stage 2 writes at least two candidates; after hand-written scout
JSON files for one extract_enum_candidate and one
introduce_fk_candidate, Stage 4 writes report.md and findings.json
whose bucket counts match the scout files.
Non-goals
- Writing the TextChoices enum proposal (that's
/extract-enum). - Writing the FK + data-migration proposal (that's
/introduce-fk). - Executing any refactor (that's
/fix-workflowafter proposal approval). - Running tests — read-only audit; tests run during the refactor skill.
- CI gates — periodic audit, not a per-commit check. The lint rules
stringly-statusandquery-mutationcover per-commit guarding. - Resolving cross-module TextChoices references — the detector skips these; the scout reads the model file to confirm.
When things go sideways
| Symptom | Action |
|---|---|
| Stage 1 detect.py finds 0 hits | Target has no implicit state (best outcome) — or the directory argument is wrong. Re-run with a valid --target <dir> |
| Any script exits non-zero | Stop at the failing stage, paste the exact command and stderr, and do not summarize downstream artifacts from a previous run |
| Stage 1 detect.py is slow (>2 min) | Shouldn't happen on a typical source tree — check for __pycache__ entries in target. Add more --skip-file-glob flags if needed |
| Stage 2 reports 0 candidates | Same as Stage 1 zero — or collapse ignored all hits (check stderr) |
Stage 3 scout buckets everything as enum_already_used |
Scout is being too permissive. Inspect one output; re-dispatch with "consult knowledge/ for known enum list" |
Scout recommends /introduce-fk for a freshness-check hit |
Rule 1 in verify.md was skipped — re-dispatch citing freshness_not_identity |
| Report lists noqa'd candidates as actionable | Detector-only bug — the scout should have bucketed as legacy_allow_list. Re-dispatch with a # noqa grep in the prompt |
Repository layout
.claude/skills/find-implicit-state/
├── SKILL.md # this file — orchestrator
├── scripts/
│ ├── detect.py # Stage 1
│ ├── collapse.py # Stage 2
│ └── report.py # Stage 4
├── agents/
│ └── verify.md # Stage 3 scout brief
└── knowledge/ # sub-agent context, never loaded by orchestrator
└── verification.md
The orchestrator (you) never reads files in knowledge/. Those
are for the scout sub-agents. Keeping them out of your context is the
whole point of this architecture.