Pipeline Auditor (draft audit + regression)
Purpose: a deterministic “regression test” for the writing stage.
It answers:
- did we leak placeholders or planner talk?
- did citation scope drift?
- did the draft fall back to generator voice (navigation/narration templates)?
- is citation density/health sufficient for a survey-like draft?
This skill is analysis-only. It does not edit content. For survey/deep, style/citation-shape violations are blocking by default.
Inputs
output/DRAFT.md
outline/outline.yml
- Optional (recommended):
outline/evidence_bindings.jsonl
citations/ref.bib
Outputs
What it checks (deterministic)
A150++ citation targets (used by the auditor):
Per-H3: >=12 unique citations (deep: >=14).
Global: >=150 unique citations across the full draft (recommended target: 165; deep floor: 165).
Placeholder leakage: ellipsis (..., …), TODO markers, scaffold tags.
Outline alignment: section/subsection order vs outline/outline.yml.
Survey tables (survey deliverable): require >=2 Markdown tables in the merged draft (index tables live in outline/tables_index.md) (inserted by section-merger from outline/tables_appendix.md).
Paper voice anti-patterns:
- narration templates (
This subsection ..., In this subsection ...)
- slide navigation (
Next, we move ..., We now turn to ...)
- pipeline voice (
this run, “pipeline/stage/workspace” in prose)
Evidence-policy disclaimer spam: repeated “abstract-only/title-only/provisional” boilerplate inside H3 bodies.
Meta survey-guidance phrasing: survey synthesis/comparisons should ....
Synthesis stem repetition: repeated Taken together, ... and similar high-signal generator stems.
Numeric claim context: numbers without minimal evaluation context tokens (benchmark/dataset/metric/budget/cost).
Citation health (if citations/ref.bib exists): undefined keys, duplicates, basic formatting red flags.
Citation-shape hard gate (survey/deep): no adjacent citation blocks ([@a] [@b]), no duplicate keys inside one block ([@a; @a]), and per-H3 mid-sentence citation ratio >=30%.
Citation scope (if outline/evidence_bindings.jsonl exists): citations used per H3 should stay within the bound evidence set.
How to use the report (routing table)
Treat output/AUDIT_REPORT.md as a “what to fix next” router.
Common FAIL families -> responsible stage/skill:
Placeholders / leaked scaffolds
- Fix: C2–C4 artifacts are not clean. Route to
subsection-briefs / evidence-draft / writer-context-pack, then rewrite affected sections.
Missing overview tables (draft has <2 tables)
- Fix: ensure
table-schema + appendix-table-writer produced outline/tables_appendix.md (>=2 tables, citation-backed, no placeholders), then rerun section-merger (tables insert as an Appendix block by default).
Planner talk in transitions / narrator bridges
- Fix: rerun
transition-weaver (and ensure briefs include bridge_terms / contrast_hook), then re-merge.
Narration templates / slide navigation inside H3
- Fix: rewrite the failing
sections/S*.md via writer-selfloop (local, section-level) or subsection-polisher.
Evidence-policy disclaimer spam
- Fix: keep evidence policy once in Intro/Related Work (front matter), delete repeats in H3 (use
draft-polisher or local section rewrites).
Citation scope drift (out-of-scope bibkeys)
- Fix: either (a) rewrite the subsection to stay in-scope, or (b) fix mapping/bindings (
section-mapper → evidence-binder) and regenerate packs.
Global unique citations too low
- Fix:
citation-diversifier → citation-injector (NO NEW FACTS), then draft-polisher.
Intro/Related Work too thin / too few cites
- Fix: rewrite the corresponding
sections/S<sec_id>.md front-matter file via writer-selfloop (front-matter path) using dense positioning + method paragraph.
Prevention guidance (what upstream writers should do)
If you want the auditor to PASS without a heavy polish loop:
- Start each H3 with a content claim + thesis (avoid narration templates).
- Use explicit contrasts and at least one evaluation anchor paragraph.
- Embed citations per claim (avoid trailing cite dumps).
- Put evidence-policy limitations once in the front matter, not in every H3.
Script
Quick Start
python .codex/skills/pipeline-auditor/scripts/run.py --help
python .codex/skills/pipeline-auditor/scripts/run.py --workspace workspaces/<ws>
All Options
--workspace <dir>
--unit-id <U###> (optional; for logs)
--inputs <semicolon-separated> (rare override; prefer defaults)
--outputs <semicolon-separated> (rare override; default writes output/AUDIT_REPORT.md)
--checkpoint <C#> (optional)
Examples
- Run audit after
global-reviewer and before LaTeX/PDF:
python .codex/skills/pipeline-auditor/scripts/run.py --workspace workspaces/<ws>
Troubleshooting
Issue: audit fails due to undefined citations
Fix:
- Regenerate citations with
citation-verifier and ensure citations/ref.bib contains every cited key.
Issue: audit fails due to narration-style navigation phrases
Fix:
- Rewrite as argument bridges (content-bearing handoffs, no navigation commentary) in the failing
sections/* files, then re-merge.
Issue: audit fails due to "unique citations too low"
Fix:
- Run
citation-diversifier to produce output/CITATION_BUDGET_REPORT.md.
- Apply it via
citation-injector (edits output/DRAFT.md, writes output/CITATION_INJECTION_REPORT.md).
- Then run
draft-polisher → global-reviewer → auditor.
1---2name: pipeline-auditor3description: Audit/regression checks for the evidence-first survey pipeline: citation health, per-section coverage, placeholder leakage, and template repetition. **Trigger**: auditor, audit, regression test, quality report, 审计, 回归测试. **Use when**: `output/DRAFT.md` exists and you want a deterministic PASS/FAIL report before LaTeX/PDF. **Skip if**: you are still changing retrieval/outline/evidence packs heavily (audit later). **Network**: none. **Guardrail**: do not change content; only analyze and report.4---5
6# Pipeline Auditor (draft audit + regression)
7
8Purpose: a deterministic “regression test” for the writing stage.
9
10It answers:
11- did we leak placeholders or planner talk?
12- did citation scope drift?
13- did the draft fall back to generator voice (navigation/narration templates)?
14- is citation density/health sufficient for a survey-like draft?
15
16This skill is analysis-only. It does not edit content. For `survey`/`deep`, style/citation-shape violations are blocking by default.
17
18## Inputs
19
20- `output/DRAFT.md`
21- `outline/outline.yml`
22- Optional (recommended):
23 - `outline/evidence_bindings.jsonl`
24 - `citations/ref.bib`
25
26## Outputs
27
28- `output/AUDIT_REPORT.md`
29
30## What it checks (deterministic)
31
32A150++ citation targets (used by the auditor):
33- Per-H3: >=12 unique citations (deep: >=14).
34- Global: >=150 unique citations across the full draft (recommended target: 165; deep floor: 165).
35
36- Placeholder leakage: ellipsis (`...`, `…`), TODO markers, scaffold tags.
37- Outline alignment: section/subsection order vs `outline/outline.yml`.
38- Survey tables (survey deliverable): require >=2 Markdown tables in the merged draft (index tables live in `outline/tables_index.md`) (inserted by `section-merger` from `outline/tables_appendix.md`).
39- Paper voice anti-patterns:
40 - narration templates (`This subsection ...`, `In this subsection ...`)
41 - slide navigation (`Next, we move ...`, `We now turn to ...`)
42 - pipeline voice (`this run`, “pipeline/stage/workspace” in prose)
43- Evidence-policy disclaimer spam: repeated “abstract-only/title-only/provisional” boilerplate inside H3 bodies.
44- Meta survey-guidance phrasing: `survey synthesis/comparisons should ...`.
45- Synthesis stem repetition: repeated `Taken together, ...` and similar high-signal generator stems.
46- Numeric claim context: numbers without minimal evaluation context tokens (benchmark/dataset/metric/budget/cost).
47- Citation health (if `citations/ref.bib` exists): undefined keys, duplicates, basic formatting red flags.
48- Citation-shape hard gate (`survey`/`deep`): no adjacent citation blocks (`[@a] [@b]`), no duplicate keys inside one block (`[@a; @a]`), and per-H3 mid-sentence citation ratio >=30%.
49- Citation scope (if `outline/evidence_bindings.jsonl` exists): citations used per H3 should stay within the bound evidence set.
50
51## How to use the report (routing table)
52
53Treat `output/AUDIT_REPORT.md` as a “what to fix next” router.
54
55Common FAIL families -> responsible stage/skill:
56
57- Placeholders / leaked scaffolds
58 - Fix: C2–C4 artifacts are not clean. Route to `subsection-briefs` / `evidence-draft` / `writer-context-pack`, then rewrite affected sections.
59
60- Missing overview tables (draft has <2 tables)
61 - Fix: ensure `table-schema` + `appendix-table-writer` produced `outline/tables_appendix.md` (>=2 tables, citation-backed, no placeholders), then rerun `section-merger` (tables insert as an Appendix block by default).
62
63- Planner talk in transitions / narrator bridges
64 - Fix: rerun `transition-weaver` (and ensure briefs include `bridge_terms` / `contrast_hook`), then re-merge.
65
66- Narration templates / slide navigation inside H3
67 - Fix: rewrite the failing `sections/S*.md` via `writer-selfloop` (local, section-level) or `subsection-polisher`.
68
69- Evidence-policy disclaimer spam
70 - Fix: keep evidence policy once in Intro/Related Work (front matter), delete repeats in H3 (use `draft-polisher` or local section rewrites).
71
72- Citation scope drift (out-of-scope bibkeys)
73 - Fix: either (a) rewrite the subsection to stay in-scope, or (b) fix mapping/bindings (`section-mapper` → `evidence-binder`) and regenerate packs.
74
75- Global unique citations too low
76 - Fix: `citation-diversifier` → `citation-injector` (NO NEW FACTS), then `draft-polisher`.
77
78- Intro/Related Work too thin / too few cites
79 - Fix: rewrite the corresponding `sections/S<sec_id>.md` front-matter file via `writer-selfloop` (front-matter path) using dense positioning + method paragraph.
80
81## Prevention guidance (what upstream writers should do)
82
83If you want the auditor to PASS *without* a heavy polish loop:
84- Start each H3 with a content claim + thesis (avoid narration templates).
85- Use explicit contrasts and at least one evaluation anchor paragraph.
86- Embed citations per claim (avoid trailing cite dumps).
87- Put evidence-policy limitations once in the front matter, not in every H3.
88
89## Script
90
91### Quick Start
92
93- `python .codex/skills/pipeline-auditor/scripts/run.py --help`
94- `python .codex/skills/pipeline-auditor/scripts/run.py --workspace workspaces/<ws>`
95
96### All Options
97
98- `--workspace <dir>`
99- `--unit-id <U###>` (optional; for logs)
100- `--inputs <semicolon-separated>` (rare override; prefer defaults)
101- `--outputs <semicolon-separated>` (rare override; default writes `output/AUDIT_REPORT.md`)
102- `--checkpoint <C#>` (optional)
103
104### Examples
105
106- Run audit after `global-reviewer` and before LaTeX/PDF:
107 - `python .codex/skills/pipeline-auditor/scripts/run.py --workspace workspaces/<ws>`
108
109## Troubleshooting
110
111### Issue: audit fails due to undefined citations
112
113Fix:
114- Regenerate citations with `citation-verifier` and ensure `citations/ref.bib` contains every cited key.
115
116### Issue: audit fails due to narration-style navigation phrases
117
118Fix:
119- Rewrite as argument bridges (content-bearing handoffs, no navigation commentary) in the failing `sections/*` files, then re-merge.
120
121### Issue: audit fails due to "unique citations too low"
122
123Fix:
124- Run `citation-diversifier` to produce `output/CITATION_BUDGET_REPORT.md`.
125- Apply it via `citation-injector` (edits `output/DRAFT.md`, writes `output/CITATION_INJECTION_REPORT.md`).
126- Then run `draft-polisher` → `global-reviewer` → auditor.