# Claim Fit

> Claim Fit (citation-precision audit of normative-attribution claims)

- Skill: `jposluns/claim-fit` (Agent Skill)
- Install (CLI): `npx skillmds@latest add jposluns/claim-fit`
- Raw SKILL.md: https://api.skillmd.com/api/skills/jposluns/claim-fit/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Security
- Author: jposluns (https://skillmd.com/u/jposluns)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/jposluns/claim-fit

---


# Claim Fit (citation-precision audit of normative-attribution claims)

## Project wiring (the parent library's instantiation; adopters substitute their own)

Portable procedure, concrete names. In the parent GRC library this skill runs with:

- Citation, existence, and currency gates (the mechanical floor this skill's precision layer sits on): gates 5 and 6 (the framework-citation hallucination audit and the standards-currency audit), plus the control-code existence gates 48, 49, 54, 58, and 61 for cited control identifiers.
- Recall-oriented pre-filter: `tools/audit-claim-precision.py` (advisory, always exits 0, explicitly not a gate), with `--ref-base <path-to-reference-base-checkout>` to report each named source family's held-state from the reference-base indexes and `--tier A` for the Tier-A-only cadence.
- Mechanical baseline: `tools/run_all_audits.sh` (the all-gates suite step 1 confirms exits 0 before the semantic read).
- Reference base: the sibling private `grc_library_ref` repository, located via its indexes (`grc_library_ref/INDEX.md`, `grc_library_ref/catalogue.yml`, `grc_library_ref/SECTION-INDEX.md`), with the per-source currency confirmation and trust-bucket rules its own conventions define.
- Tier definitions (as the pre-filter implements them): TIER A, a specific value tied to a named normative source in one clause, both orders, judged EVERY row; TIER B, soft-alignment phrasing with no specific value, sampled on cadence. RESIDUE: the pre-filter keys Tier A on a DIGIT-bearing value, so a value in WORD FORM (e.g. "reviewed at least annually") is captured as Tier B rather than judge-every-row Tier A; "every Tier-A row judged" is therefore read at that digit-keyed width (a word-form value-attribution is sampled with the rest of Tier B, not judged every run).

- QA execution (this project): the pass itself runs as the TRIPLE-FAMILY QA panel (the CLAUDE.md standard), the identical brief given to one Claude-, one Codex-, and one Gemini-family orch-verify worker with the verdicts reconciled by UNION (any family's non-`prescribed` candidate for a (value, source) pair enters step 4 and the orchestrator adjudicates it against the held text; a 2-1 split is NEVER resolved by majority-discard); the Process's "dispatch one or more subagents (or perform the read directly)" wording describes the WORKER's own internal execution, never an in-session Task/Agent dispatch by the maintainer orchestrator, which the offload rule forbids and the block-orchestrator-self-qa hook blocks.
- Held-source citation and holdings/currency basis (#1926, from the 2026-08-31 stale-note `grc_library_ref` incident): every verdict cites the HELD SOURCE's own reference-base `path:line` (where the passage it read is located) as a REQUIRED field alongside the quoted passage (or, for `source-not-held`, the executed index lookup that returned nothing), and a verdict without that evidence is REJECTED, because a passage recalled from a note or from memory has no reference-base location to cite. That location comes from OPENING the held file, located via an EXECUTED `python3 tools/ref-holds.py <source>` plus a reference-base INDEX read performed in THIS run (ref-holds and the index locate the held FILE and are the authority for what is HELD; the `path:line` is where the passage sits, read within that file this run), never a `pending-decisions` note, a prior run's record, or memory. CURRENCY is a SEPARATE question: whether a held source is the CURRENT version is validated UPSTREAM this run per the reference-version-currency SOP, never inferred from the index (which is believed-current STORAGE, not a version authority).

An adopting project maps each bullet to its own gates, pre-filter, baseline runner, reference base, and tiers; the procedure below refers to them generically.

## Overview

The audit gates that guard citations check existence, well-formedness, and currency, not precision. The citation family confirms a named standard exists with the right year and identifier; the control-code gates confirm a cited code is in its catalogue. None of them asks the question that decides whether a normative-attribution sentence is true: does the cited source actually contain the supporting language the sentence attributes to it? A claim can name a real source, cite it in perfect form, and still attribute to it a value the source never states. That class, "attributed value, silent source", is gate-blind by construction, because precision cannot be decided by a string-and-membership check; it needs a reader to compare the attributed value against the source's own clause.

The class is not hypothetical. The motivating incident in the parent library: a policy attributed a fixed "180-day baseline" to NIST SP 800-53 CA-6 and ISO/IEC 27001 Clause 9.2, and neither source prescribes a fixed interval. Two further illustrative shapes from the same provenance: three carriers attributing a 7-year retention to ISO/IEC 42001 plus EU AI Act Annex IV (the ISO half informs the figure without prescribing it; the Annex IV half carries no retention obligation at all; the value was the project's own canonical choice, to be recorded as such wherever the adopting project records its decisions), and a 24-hour supplier-notification KPI citing GDPR Article 33(2), which sets "without undue delay", not a fixed clock.

`claim-fit` is the semantic-judge half of a two-part instrument whose recall-oriented triage half is the advisory pre-filter named in the project wiring above (explicitly NOT a gate; always exits 0; CI cannot host the check because the ground truth lives in the held reference base the ordinary mechanical gate cannot see). The tool tiers the worklist by risk: TIER A rows (a specific value tied to a named normative source in one clause, both orders, the exact motivating-incident shape) are judged EVERY one; TIER B rows (soft-alignment phrasing with no specific value) are sampled on cadence. The skill judges: for each worklisted claim, it reads EACH cited source's own text in the reference base and decides, PER (value, source) PAIR, whether that source prescribes, merely informs, or contradicts the attributed value. The binding rule mirrors the matrix-fit lesson: judge against the held source TEXT, never a remembered meaning, because a remembered meaning is exactly what produced the misattribution.

The verdict vocabulary is four-valued, because the right fix differs by verdict:

- **`prescribed`**: the source states the attributed value; the claim fits as written.
- **`informed-not-prescribed`**: the source motivates the value but does not state it (the 7-year retention shape illustrated above). The fix is the ATTRIBUTION PHRASING, never the value: reword to name the value as the project's or organization's canonical choice informed by the source, so the sentence stops asserting source language that does not exist.
- **`mis-attributed`**: the source states a different value, or none, where the sentence asserts one (the motivating-incident and Article 33(2) shapes). The fix is the value or the citation, decided per finding.
- **`source-not-held`**: the reference base holds no text for the named source, so no judgement is possible. The finding FIRST triggers an acquisition attempt via the project's ingest path (maintainer-side where workers are egress-constrained); only a failed or egress-constrained attempt produces a source-acquisition queue entry, which then carries named options (the maintainer provides the source, the task defers and routes around it, or the claim is reworded so it no longer depends on the unheld source). The claim is NEVER adjudicated from memory: the evidence-grounded-completion missing-load-bearing-reference rule binds (a source is acquired or the work pauses, never worked around), and routing to the queue WITHOUT first attempting the acquisition is exactly the shortcut it forecloses.

This skill is a single-pass advisory audit, not a fix-to-fixed-point loop and not a trust-recovery escalation. It runs on a cadence, surfaces confirmed misattributions, and routes or fixes them under the normal in-window / out-of-window triage. It is to normative-value claims what `/matrix-fit` is to control-code citations: the semantic layer over gates that can only check existence.

## When to Use

- **The one-time full Tier-A pass** when the skill is adopted (the whole Tier-A census is small by construction; every row is judged, each cited source assessed, establishing the baseline).
- **After any batch that adds or edits normative-value claims** (a substantive content batch, a jurisdiction annex, a KPI or SLA table): run the tool, judge the NEW Tier-A rows the batch introduced, and sample the batch's OWN Tier-B rows (this cadence is batch-scoped via `--docs`; it does not advance the global Tier-B coverage cycle below).
- **Periodically, the Tier-B coverage sweep** (the global pass): the whole Tier-B population is too large to judge every run, so a standing cadence works through it a sample at a time until the entire census has been judged once, then restarts. This is the cadence the sampling discipline in step 3 governs, distinct from the batch-scoped per-batch cadence above.
- **Ad-hoc when a claim is in doubt** (a maintainer flag, a `/validate` or `/full-qa` note about a suspicious attribution, an apply-time uncertainty about whether a source states a value).
- **NOT as a replacement for the citation gates.** The existence, currency, and control-code gates still run on every PR; `claim-fit` is the precision layer on top of them. A claim must pass the gates first; this skill judges precision among claims that already pass.

## Process

### 1. Establish scope and confirm the reference base

Name the scope for this run: the full Tier-A census (the adoption pass), the rows a just-applied batch introduced (the per-batch cadence), a flagged set (the ad-hoc cadence), or the whole Tier-B population for the periodic coverage sweep (run globally, no `--docs`). Confirm the project's all-gates suite (named in the project wiring above) exits 0 first; a precision pass judges among claims that already pass the mechanical gates. Confirm the reference base is available: its indexes locate the held source texts, and the per-source currency rule applies (confirm a held source is current before relying on it, per the reference base's own invariants; a superseded held text is grounds to defer, not to judge against the stale text silently). Every verdict cites the held source's own reference-base `path:line` (or, for `source-not-held`, the executed index lookup that returned nothing), read from the held file this run (the file located via an EXECUTED `ref-holds` + index read; the Project-wiring held-source-citation requirement), never a `pending-decisions` note, a prior run's record, or memory; a source's CURRENCY is validated upstream this run, not inferred from the index.

### 2. Run the advisory triage tool to generate the worklist

Run the recall-oriented pre-filter named in the project wiring, pointing it at the reference base (restrict to Tier A for the Tier-A-only cadence; scope to a batch's documents for the per-batch cadence; run globally with no `--docs` for the periodic Tier-B coverage sweep, so Tier B is sampled across the whole census). The tool always exits 0; its output is a recall-oriented worklist, tiered, with each named source family's held-state reported best-effort from the reference-base indexes. Treat the worklist as the judge's input-narrowing step, NOT a defect list: a listed row is a candidate to read, and a lexical extractor deliberately over-collects (a spurious row costs the judge seconds; a missed row ships a misattribution). Add any claim the maintainer or a prior QA note flagged even if the extractor did not list it.

### 3. Dispatch the citation-precision judge over the worklist

Dispatch one or more subagents (or perform the read directly for a small worklist) to judge each claim. The judge brief: locate the named source's held text via the reference-base indexes, read the specific clause, article, or annex the claim points at, and return, FOR EACH (attributed value, named source) PAIR on the claim, one of the four verdicts (`prescribed` / `informed-not-prescribed` / `mis-attributed` / `source-not-held`) with the source passage QUOTED as evidence. A claim that cites SEVERAL sources for one value is judged once PER source: the same value can draw a DIFFERENT verdict per source (e.g. `prescribed` by one but `informed-not-prescribed`, `mis-attributed`, or `source-not-held` for another). The Overview's 7-year example: `(7 years, ISO/IEC 42001)` is `informed-not-prescribed` (the ISO clause informs the figure without prescribing a period), while `(7 years, EU AI Act Annex IV)` is `mis-attributed` (Annex IV is held but states no such retention obligation). A single per-claim verdict would erase that distinction. The binding rules: judge against the held source text, never memory or plausibility; a verdict without a quoted source passage (or, for `source-not-held`, without the index lookup that failed) is a hypothesis, not a finding; and an un-held source is never adjudicated from recall, whatever the judge's confidence. Every judgement (one row per (value, source) pair) quotes the claim's location as `path:line`, the attributed value, the named source, the source passage, AND the held source's own reference-base `path:line` where that passage was read in THIS run (or, for `source-not-held`, the executed index lookup that returned nothing); a verdict without EITHER the held-source `path:line` OR, for `source-not-held`, the executed index lookup that returned nothing is REJECTED, because a passage from a note or from memory has no reference-base evidence to cite.

**Tier-B sampling (the periodic coverage sweep, step-3 discipline).** In-scope Tier-A rows are judged exhaustively every run; the Tier-B population is too large to judge in full each run (the tool reports the live count), so the Tier-B COVERAGE SWEEP cadence (When to Use) works through it a sample at a time. Each sweep run samples a default N rows (10 is proportionate to the current census; the FINAL run of a cycle samples min(N, the remaining un-judged eligible rows), so it may draw fewer than N), STRATIFIED by the JOINT (cited source family x document domain) cell so the sample spreads across families and domains rather than clustering; when the number of non-empty strata exceeds N a run draws from a rotating subset of strata and successive runs rotate the remainder, so every stratum is reached OVER THE CADENCE (a single N-row run necessarily covers only some strata, so the no-family-or-domain-skipped property is cumulative across the cadence, never per-run). A sweep run draws WITHOUT REPLACEMENT against the Tier-B rows earlier sweep runs already judged: the history record for a coverage run carries the CYCLE number, the rows it sampled (keyed by document path plus cited source plus a short claim-text anchor, NOT a bare line number, since edits shift lines), and a RESET marker written when the cycle completes and sampling restarts on the refreshed census. RESIDUE, stated at the point of use: this is an operator-tracked discipline over the history record, and its row identity is approximate (a document rename or a reworded claim still needs manual reconciliation); the history model's cross-run mechanics (the source-family and document-domain classification, the rotation cursor, and the without-replacement set-diff) are MECHANIZED by the pre-filter tool's `--sample` / `--unjudged` mode (shipped): it reads a JSONL sweep ledger, computes the un-judged set, and emits the next stratified N deterministically (no seed, date, or randomness in the draw). ONE approximation persists BY DESIGN: the row anchor is a hash of the normalized claim text, so a genuine reword changes the anchor and needs manual reconciliation. The coverage sweep is run through that mode, not hand-tracked. The per-batch and ad-hoc cadences are unaffected: they sample only their own scoped Tier-B rows and do not advance the coverage cycle. TWO further coverage residues are DISCLOSED (review, not yet mechanized): (a) the census counts a MULTI-SOURCE line once, keyed to its FIRST cited source, so "the entire census judged once" is a census of lines-by-first-source, not of every (value, source) pair, and a family habitually cited second is under-represented in the strata (the judge-side per-pair expansion still recovers verdicts for a listed line; only the coverage accounting runs on first-source keys); (b) a drawn row the judge could NOT adjudicate this run (`source-not-held` or deferred) is still recorded in the sampled list and counts as covered for the cycle, so "covered" means "drawn", not "judged". The two differ in DIRECTION: (a) UNDER-counts the census (a second-cited pair never enters, so coverage is understated, the safe direction), but (b) can OVER-state judged coverage, a drawn row the state-reader treats as judged (audit-claim-precision.py counts every sampled key as covered) is dropped from later draws even if it was never adjudicated, so it may go permanently unjudged, the consequential direction the per-row judged-flag mechanization exists to close. Both are queued for the finditer-census / judged-flag mechanization; until then (b) is the one to watch, so a sweep run NAMES any drawn row it could not adjudicate rather than letting the ledger imply it was.

### 4. Synthesize and apply-time-verify each candidate against the source text

The orchestrator re-reads each candidate `mis-attributed` and `informed-not-prescribed` verdict's source passage in the reference base before treating it as a finding (the judge produces research; the orchestrator confirms). A judge false positive (a passage the judge missed elsewhere in the source, a version mismatch, an over-literal reading of a clause the source states in different words) is refuted here, not routed. For each confirmed finding, draft the fix per the verdict; for a MULTI-SOURCE claim the composite fix derives from the VECTOR of per-pair verdicts (e.g. prescribed-by-A + silent-in-B: attribute the value to A alone, or split the attribution so each source carries only what it prescribes): a rephrased attribution for `informed-not-prescribed` (the value stands; the sentence stops asserting source language), a corrected value or citation for `mis-attributed` (which one is a per-finding call, surfaced to the maintainer when the value is an authorial choice), and, for `source-not-held`, an acquisition attempt FIRST (maintainer-side where egress-constrained), producing a source-acquisition queue entry only on a failed or constrained attempt, per the verdict definition above. When an attribution is rephrased, grep the full touched file AND the corpus in-scope content for sibling carriers of the same attribution at bare-token width (the same value-plus-source pair recurs across policy families, not just the touched file); `check-class-completeness.py` is the named instrument.

### 5. Triage and route findings

For confirmed findings in the current scope, fix them in-window: apply the fix, bump the touched document's version and date metadata in the same commit, and record the correction in the project's detailed change record. Where the fix requires an authorial decision (which of value-vs-citation is wrong; whether a canonical project value should be re-anchored), surface it to the maintainer with named options rather than silently picking. Confirmed findings outside the current scope are surfaced with named options (fix-now vs route-to-backlog), not auto-deferred. Findings refuted at apply-time are recorded with the refutation, not routed. Findings that dedupe against an existing backlog item are cross-referenced, not duplicated.

### 6. Record and surface

Surface confirmed findings inline in chat (per finding, one per (value, source) pair: claim `path:line`, the attributed value, the named source, the verdict with the quoted source passage and its held-source reference-base `path:line` (or, for `source-not-held`, the executed index lookup that returned nothing), the fix applied or the option surfaced). Write the run to the project's claim-fit record location and append a history row; a zero-finding run still gets a history row (the proof-of-discipline), with no detail file. The pass terminates when the worklist is judged, the confirmed findings are routed or fixed, and the run is recorded; it is a single advisory pass, not a fix-to-fixed-point loop.

## Red Flags

- Judging a claim from the source's remembered content or general reputation instead of reading the held text. The remembered meaning is the failure mode that produced the misattribution.
- Adjudicating a `source-not-held` claim anyway "because the answer is well known". No held text, no judgement; the claim triggers an acquisition attempt FIRST, and routes to the source-acquisition queue only on a failed or egress-constrained attempt (per the verdict definition).
- Recording a verdict, or a held/not-held holdings claim, from a `pending-decisions` note or a prior run's record instead of an executed `ref-holds` + index read (the held-source `path:line`, or the failed index lookup for `source-not-held`) cited THIS run (the 2026-08-31 stale-note incident this cadence's held-source-citation requirement exists to prevent).
- Treating the triage tool's worklist as a defect list. It is recall-oriented; most Tier-A rows in clean project content will judge `prescribed`, and Tier-B rows are soft-alignment claims that mostly fit.
- Fixing an `informed-not-prescribed` finding by changing the VALUE. The value is frequently the project's own canonical choice (the canonical-choice class illustrated above); the defect is the attribution phrasing, and silently changing a canonical value is an authorial decision the maintainer owns.
- Routing a judge verdict without the orchestrator's own re-read of the source passage. Apply-time verification is the false-positive filter; a judge can miss the prescribing clause elsewhere in a long source.
- Rephrasing one carrier of a misattributed value while sibling carriers of the same value-plus-source pair survive. Grep at bare-token width across the touched file and the project's in-scope content before claiming the class fixed.
- Running this as a substitute for the citation gates, or skipping it because "the gates passed". The gates and this skill cover orthogonal classes; a well-formed citation says nothing about precision.

## Verification

The pass is complete on a given run when:

- The scope was named and the mechanical baseline was clean (the project's all-gates suite exit 0) before the semantic read.
- The triage tool was run and its worklist (plus any flagged claims) was the judge's input.
- Every in-scope Tier-A row was judged against the held source text, EACH (value, source) pair on a multi-source row assessed, with the passage quoted (or verdicted `source-not-held` on index evidence). For a Tier-B COVERAGE-SWEEP run, the sample was drawn per the step-3 discipline (default N, joint source-family x domain strata rotated across runs) and NAMED in the history record with the cycle number, the sampled rows (keyed by path + cited source + claim-text anchor), and any reset marker, so a later run can draw without replacement against it; for a per-batch or ad-hoc run, the Tier-B rows sampled within that scope are named.
- Every verdict cited the held source's own reference-base `path:line` from an executed `ref-holds` + index read in THIS run (not a `pending-decisions` note, a prior run's record, or memory); a `source-not-held` verdict cited the executed index lookup that returned nothing, and any currency claim was validated upstream this run rather than inferred from the index.
- The orchestrator re-read each candidate finding's source passage and refuted or confirmed it; refutations are recorded, not routed.
- Confirmed in-scope findings were fixed per their verdict class (version and date metadata bumped, change-record entry written) or surfaced with named options where the fix is authorial; `source-not-held` claims triggered an acquisition attempt first and were queued only on a failed or egress-constrained attempt.
- Any in-scope row not judged (deferred, routed, or `source-not-held`) was NAMED with its blocker in the record; none was silently dropped.
- The run was recorded (history row always; detail file when findings exist) and findings were surfaced inline in chat.

## Common Rationalizations

| Rationalization | Reality |
|---|---|
| "The citation gates pass, so the claim is fine." | The gates check the source exists and the citation is well-formed, not that the source states the attributed value. Only a read of the clause decides. |
| "Everyone knows GDPR Article 33 says 72 hours." | Article 33(1) does; the flagged carrier among the illustrative shapes cites 33(2), which says "without undue delay". The clause read, not the reputation, is the evidence. |
| "The source is not held, but I am confident what it says." | An un-held source is never adjudicated from memory (the external-version corollary). Attempt acquisition FIRST; route to the source-acquisition queue only on a failed or egress-constrained attempt. |
| "The reference base held this last run, or a note says it is held." | Holdings is determined by an executed `ref-holds` + index read THIS run, cited by the held-source `path:line` (or the failed index lookup for a not-held source); a note or a prior run's record is not a holdings authority, and a source's currency is validated upstream, never inferred from the index. |
| "The source does not state the value, so the value is wrong." | Often the value is the project's own canonical choice the source merely informs. The fix is the attribution phrasing; changing the value is the maintainer's call. |
| "The worklist is short, so the project content is precise." | The worklist is recall-oriented triage over lexical shapes; unlisted prose can still misattribute in shapes the extractor does not match. A short worklist narrows the read; it certifies nothing. |
| "Claim precision should just be a gate." | It is not mechanically checkable, and the ground truth lives in a held reference base the ordinary mechanical gate cannot see. The cadenced audit is the durable instrument (the same conclusion the matrix-fit design reached). |

## See Also

- Canonical rule [`evidence-grounded-completion`](../../governance/evidence-grounded-completion.md): the assertion-side discipline this skill applies to normative attributions (a claim that a source states a value requires reading the source, not inferring it), including the external-version-currency corollary that forbids judging un-held or unconfirmed sources from memory.
- Related skill [`matrix-fit`](../matrix-fit/SKILL.md) (`/matrix-fit`): the sibling semantic audit for control-code fit; this skill applies the same pattern (advisory recall-oriented tool plus cadenced semantic judge) to attributed values. The two-PR shipping precedent and the judge-against-the-source rule both come from it.
- Related skill [`citation-quote-verification`](../citation-quote-verification/SKILL.md): verifies cited *quotes* match source text verbatim; this skill verifies attributed *values and requirements* are actually prescribed by the source. A quote can be verbatim while the surrounding attribution overstates it.
- Related skill [`validation-sweep`](../validation-sweep/SKILL.md) (`/validate`): the project-wide drift sweep whose notes can flag an attribution doubt for this skill to adjudicate.
- Related skill [`publication-screening`](../publication-screening/SKILL.md) (`/screen-publications`): the admission-control screen for configured untrusted sources; once a screened source's claim enters project content, its precision is this skill's cadence to adjudicate like any other attributed value.
- The recall-oriented pre-filter named in the project wiring above: the triage step that feeds this skill's worklist (not a gate; always exits 0; a Tier-A-only mode for the judge-every-row tier; a reference-base flag to report source held-state from the reference-base indexes).
- The reference base named in the project wiring above: located via its own indexes, with the per-source currency confirmation and trust-bucket rules its own conventions define.

