Draft Polisher (Audit-style editing)
Triggers & routing
- Trigger: polish draft, de-template, coherence pass, remove boilerplate, 润色, 去套话, 去重复, 统一术语.
- Use when: a first-pass draft exists but reads like scaffolding (repetition/ellipsis/template phrases) or needs a coherence pass before global review/LaTeX.
Goal: turn a first-pass draft into readable survey prose without breaking the evidence contract.
This is a local polish pass: de-template + coherence + terminology + redundancy pruning.
Note: if the main issue is structural redundancy from section accumulation, push the change upstream to sections/ and use paragraph-curator before merge. draft-polisher should not be the primary place where you decide which paragraphs to keep.
Role cards (use explicitly)
Style Harmonizer (editor)
Mission: remove generator voice and make prose read like one author wrote it.
Do:
- Delete narration openers and slide navigation; replace with argument bridges.
- Vary rhythm; remove repeated template stems.
- Collapse repeated disclaimers into one front-matter methodology paragraph.
Avoid:
- Adding or removing citation keys.
- Moving citations across subsections.
Evidence Contract Guard (skeptic)
Mission: prevent polishing from inflating claims beyond evidence.
Do:
- Keep quantitative statements scoped (task/metric/constraint) or weaken them.
- Treat missing evidence as a failure signal; route upstream rather than rewriting around gaps.
Avoid:
- Overconfident language when evidence is abstract-only.
Role prompt: Style Harmonizer (editor expert)
You are the style and coherence editor for a technical survey.
Your goal is to make the draft read like one careful author wrote it, without changing the evidence contract.
Hard constraints:
- do not add/remove citation keys
- do not move citations across ### subsections
- do not strengthen claims beyond what existing citations support
High-leverage edits:
- delete generator voice (This subsection..., Next we move..., We now turn...)
- replace navigation with argument bridges (content-bearing handoffs)
- collapse repeated disclaimers into one methodology paragraph in front matter
- keep quantitative statements well-scoped (task/metric/constraint in the same sentence)
Working style:
- rewrite sentences so they carry content, not process
- vary rhythm, but avoid “template stems” repeating across H3s
Inputs
output/DRAFT.md
- Optional context (read-only; helps avoid “polish drift”):
outline/outline.yml
outline/subsection_briefs.jsonl
outline/evidence_drafts.jsonl
citations/ref.bib
Outputs
output/DRAFT.md (in-place refinement)
output/citation_anchors.prepolish.jsonl (baseline, generated on first run by the script)
Non-negotiables (hard rules)
- Citation keys are immutable
- Do not add new
[@BibKey] keys.
- Do not delete citation markers.
- If
citations/ref.bib exists, do not introduce any key that is not defined there.
- Citation anchoring is immutable
- Do not move citations across
### subsections.
- If you must restructure across subsections, stop and push the change upstream (outline/briefs/evidence), then regenerate.
- No evidence inflation
- If a sentence sounds stronger than the evidence level (abstract-only), rewrite it into a qualified statement.
- When in doubt, check the subsection’s evidence pack in
outline/evidence_drafts.jsonl and keep claims aligned to snippets.
- Citation shape normalization
- Merge adjacent citation blocks in the same sentence (avoid
[@a] [@b]).
- Deduplicate keys inside one block (avoid
[@a; @a]).
- Avoid tail-only citation dumps: keep some citations in the claim sentence itself (mid-sentence), not only paragraph end.
- Quantitative claim hygiene
- If you keep a number, ensure the sentence also states (without guessing): task type + metric definition + relevant constraint (budget/cost/tool access), and the citation is embedded in that sentence.
- Avoid ambiguous model naming (e.g., “GPT-5”) unless the cited paper uses that exact label; otherwise use the paper’s naming or a neutral description.
- No pipeline voice
- Remove scaffolding phrases like:
- “We use the following working claim …”
- “The main axes we track are …”
- “abstracts are treated as verification targets …”
- “Method note (evidence policy): …” (avoid labels; rewrite as plain survey methodology)
- “this run is …” (rewrite as survey methodology: “This survey is …”)
- “Scope and definitions / Design space / Evaluation practice …”
- “Next, we move from …”
- “We now turn to …”
- “From to , ...” (title narration; rewrite as an argument bridge)
- “In the next section/subsection …”
- “Therefore/As a result, survey synthesis/comparisons should …” (rewrite as literature-facing observation)
- Also remove generator-like thesis openers that read like outline narration:
- “This subsection surveys …”
- “This subsection argues …”
Three passes (recommended)
Pass 1 — Subsection polish (structure + de-template)
Best-of-2 micro-polish (recommended):
- For any sentence/paragraph you touch, draft 2 candidate rewrites, then keep the better one.
- Choose with a simple rubric: move clarity, no template stem, citations stay anchored, and citation shape stays reader-facing (no adjacent cite blocks / dup keys).
- Do not keep both candidates. Pick one and move on (the goal is convergence, not endless rewriting).
Role split:
- Editor: rewrite sentences for clarity and flow.
- Skeptic: deletes any generic/template sentence.
Targets:
- Each H3 reads like: tension → contrast → evidence → limitation.
- Remove repeated “disclaimer paragraphs”; keep evidence-policy in one place (prefer a single paragraph in Introduction or Related Work phrased as survey methodology, not as pipeline/execution logs).
- Use
outline/outline.yml (if present) to avoid heading drift during edits.
- If present, use
outline/subsection_briefs.jsonl to keep each H3’s scope/RQ consistent while improving flow.
- Do a quick “pattern sweep” (semantic, not mechanical):
- delete outline narration:
This subsection ..., In this subsection ...
- delete slide navigation:
Next, we move from ..., We now turn to ..., In the next section ...
- delete title narration:
From <X> to <Y>, ...
- replace with: content claims + argument bridges + organization sentences (no new facts/citations)
- If
citation-injector was used, smooth any budget-injection sentences so they read paper-like:
- Keep the citation keys unchanged.
- Avoid list-injection stems (e.g., “A few representative references include …”, “Notable lines of work include …”, “Concrete examples ... include ...”).
- Prefer integrating the added citations into an existing argument sentence, or rewrite as a short parenthetical
e.g., ... clause tied to the subsection’s lens (no new facts).
- Vary phrasing; avoid repeating the same opener stem across many H3s.
- Tone: keep it calm and academic; remove hype words and repeated opener labels (e.g., literal
Key takeaway: across many H3s).
- Reduce repeated synthesis stems (e.g., many paragraphs starting with
Taken together, ...); vary synthesis phrasing and keep it content-bearing.
- Treat repeated "Taken together," as a generator-voice smell. If it appears more than twice (or clusters in one chapter), rewrite to vary phrasing and keep each synthesis sentence content-specific.
- Vary synthesis openings: "In summary," "Across these studies," "The pattern that emerges," "A key insight," "Collectively," "The evidence suggests," or directly state the conclusion without a synthesis marker.
- Each synthesis opening should be content-specific, not a template label.
Rewrite recipe for subsection openers (paper voice, no new facts):
- Delete:
This subsection surveys/argues... / In this subsection, we...
- Replace with a compact opener that does 2–3 of these (no labels; vary across subsections):
- Content claim: the subsection-specific tension/trade-off (optionally with 1–2 embedded citations)
- Why it matters: link the claim to evaluation/engineering constraints (benchmark/protocol/cost/tool access)
- Preview: what you will contrast next and on what lens (A vs B; then evaluation anchors; then limitations)
- Example skeletons (paraphrase; don’t reuse verbatim):
- Tension-first:
A central tension is ...; ...; we contrast ...
- Decision-first:
For builders, the crux is ...; ...
- Lens-first:
Seen through the lens of ..., ...
Pass 2 — Terminology normalization
Role split:
- Taxonomist: chooses canonical terms and synonym policy.
- Integrator: applies consistent replacements across the draft.
Targets:
- One concept = one name across sections.
- Headings, tables, and prose use the same canonical terms.
Pass 3 — Redundancy pruning (global repetition)
Role split:
- Compressor: collapses repeated boilerplate.
- Narrative keeper: ensures removing repetition does not break the argument chain.
Targets:
- Cross-section repeated intros/outros are removed.
- Only subsection-specific content remains inside subsections.
Script
Quick Start
uv run python .codex/skills/draft-polisher/scripts/run.py --help
uv run python .codex/skills/draft-polisher/scripts/run.py --workspace <workspace>
All Options
--workspace <dir>: workspace root
--unit-id <U###>: unit id (optional; for logs)
--inputs <semicolon-separated>: override inputs (rare; prefer defaults)
--outputs <semicolon-separated>: override outputs (rare; prefer defaults)
--checkpoint <C#>: checkpoint id (optional; for logs)
Examples
Acceptance checklist
Troubleshooting
Issue: polishing causes citation drift across subsections
Fix:
- Keep citations inside the same
### subsection; if restructuring is intentional, delete output/citation_anchors.prepolish.jsonl and regenerate a new baseline.
Issue: draft polishing is requested before writing approval
Fix:
- Record the relevant approval in
DECISIONS.md (typically Approve C2) before doing prose-level edits.
1---2name: draft-polisher3description: Audit-style editing pass for `output/DRAFT.md`: remove template boilerplate, improve coherence, and enforce citation anchoring.4---56# Draft Polisher (Audit-style editing)78## Triggers & routing910- **Trigger**: polish draft, de-template, coherence pass, remove boilerplate, 润色, 去套话, 去重复, 统一术语.11- **Use when**: a first-pass draft exists but reads like scaffolding (repetition/ellipsis/template phrases) or needs a coherence pass before global review/LaTeX.121314Goal: turn a first-pass draft into readable survey prose **without breaking the evidence contract**.1516This is a local polish pass: de-template + coherence + terminology + redundancy pruning.1718Note: if the main issue is structural redundancy from section accumulation, push the change upstream to `sections/` and use `paragraph-curator` before merge. `draft-polisher` should not be the primary place where you decide which paragraphs to keep.19202122## Role cards (use explicitly)2324### Style Harmonizer (editor)2526Mission: remove generator voice and make prose read like one author wrote it.2728Do:29- Delete narration openers and slide navigation; replace with argument bridges.30- Vary rhythm; remove repeated template stems.31- Collapse repeated disclaimers into one front-matter methodology paragraph.3233Avoid:34- Adding or removing citation keys.35- Moving citations across subsections.3637### Evidence Contract Guard (skeptic)3839Mission: prevent polishing from inflating claims beyond evidence.4041Do:42- Keep quantitative statements scoped (task/metric/constraint) or weaken them.43- Treat missing evidence as a failure signal; route upstream rather than rewriting around gaps.4445Avoid:46- Overconfident language when evidence is abstract-only.474849## Role prompt: Style Harmonizer (editor expert)5051```text52You are the style and coherence editor for a technical survey.5354Your goal is to make the draft read like one careful author wrote it, without changing the evidence contract.5556Hard constraints:57- do not add/remove citation keys58- do not move citations across ### subsections59- do not strengthen claims beyond what existing citations support6061High-leverage edits:62- delete generator voice (This subsection..., Next we move..., We now turn...)63- replace navigation with argument bridges (content-bearing handoffs)64- collapse repeated disclaimers into one methodology paragraph in front matter65- keep quantitative statements well-scoped (task/metric/constraint in the same sentence)6667Working style:68- rewrite sentences so they carry content, not process69- vary rhythm, but avoid “template stems” repeating across H3s70```7172## Inputs7374- `output/DRAFT.md`75- Optional context (read-only; helps avoid “polish drift”):76 - `outline/outline.yml`77 - `outline/subsection_briefs.jsonl`78 - `outline/evidence_drafts.jsonl`79 - `citations/ref.bib`8081## Outputs8283- `output/DRAFT.md` (in-place refinement)84- `output/citation_anchors.prepolish.jsonl` (baseline, generated on first run by the script)8586## Non-negotiables (hard rules)87881) **Citation keys are immutable**89- Do not add new `[@BibKey]` keys.90- Do not delete citation markers.91- If `citations/ref.bib` exists, do not introduce any key that is not defined there.92932) **Citation anchoring is immutable**94- Do not move citations across `###` subsections.95- If you must restructure across subsections, stop and push the change upstream (outline/briefs/evidence), then regenerate.96973) **No evidence inflation**98- If a sentence sounds stronger than the evidence level (abstract-only), rewrite it into a qualified statement.99- When in doubt, check the subsection’s evidence pack in `outline/evidence_drafts.jsonl` and keep claims aligned to snippets.1001014) **Citation shape normalization**102- Merge adjacent citation blocks in the same sentence (avoid `[@a] [@b]`).103- Deduplicate keys inside one block (avoid `[@a; @a]`).104- Avoid tail-only citation dumps: keep some citations in the claim sentence itself (mid-sentence), not only paragraph end.1051065) **Quantitative claim hygiene**107- If you keep a number, ensure the sentence also states (without guessing): task type + metric definition + relevant constraint (budget/cost/tool access), and the citation is embedded in that sentence.108- Avoid ambiguous model naming (e.g., “GPT-5”) unless the cited paper uses that exact label; otherwise use the paper’s naming or a neutral description.1091106) **No pipeline voice**111- Remove scaffolding phrases like:112 - “We use the following working claim …”113 - “The main axes we track are …”114 - “abstracts are treated as verification targets …”115 - “Method note (evidence policy): …” (avoid labels; rewrite as plain survey methodology)116 - “this run is …” (rewrite as survey methodology: “This survey is …”)117 - “Scope and definitions / Design space / Evaluation practice …”118 - “Next, we move from …”119 - “We now turn to …”120 - “From <X> to <Y>, ...” (title narration; rewrite as an argument bridge)121 - “In the next section/subsection …”122 - “Therefore/As a result, survey synthesis/comparisons should …” (rewrite as literature-facing observation)123- Also remove generator-like thesis openers that read like outline narration:124 - “This subsection surveys …”125 - “This subsection argues …”126127## Three passes (recommended)128129### Pass 1 — Subsection polish (structure + de-template)130131Best-of-2 micro-polish (recommended):132- For any sentence/paragraph you touch, draft 2 candidate rewrites, then keep the better one.133- Choose with a simple rubric: move clarity, no template stem, citations stay anchored, and citation shape stays reader-facing (no adjacent cite blocks / dup keys).134- Do not keep both candidates. Pick one and move on (the goal is convergence, not endless rewriting).135136Role split:137- **Editor**: rewrite sentences for clarity and flow.138- **Skeptic**: deletes any generic/template sentence.139140Targets:141- Each H3 reads like: tension → contrast → evidence → limitation.142- Remove repeated “disclaimer paragraphs”; keep evidence-policy in **one** place (prefer a single paragraph in Introduction or Related Work phrased as survey methodology, not as pipeline/execution logs).143- Use `outline/outline.yml` (if present) to avoid heading drift during edits.144- If present, use `outline/subsection_briefs.jsonl` to keep each H3’s scope/RQ consistent while improving flow.145- Do a quick “pattern sweep” (semantic, not mechanical):146 - delete outline narration: `This subsection ...`, `In this subsection ...`147 - delete slide navigation: `Next, we move from ...`, `We now turn to ...`, `In the next section ...`148 - delete title narration: `From <X> to <Y>, ...`149 - replace with: content claims + argument bridges + organization sentences (no new facts/citations)150- If `citation-injector` was used, smooth any budget-injection sentences so they read paper-like:151 - Keep the citation keys unchanged.152 - Avoid list-injection stems (e.g., “A few representative references include …”, “Notable lines of work include …”, “Concrete examples ... include ...”).153 - Prefer integrating the added citations into an existing argument sentence, or rewrite as a short parenthetical `e.g., ...` clause tied to the subsection’s lens (no new facts).154 - Vary phrasing; avoid repeating the same opener stem across many H3s.155- Tone: keep it calm and academic; remove hype words and repeated opener labels (e.g., literal `Key takeaway:` across many H3s).156- **Reduce repeated synthesis stems** (e.g., many paragraphs starting with `Taken together, ...`); vary synthesis phrasing and keep it content-bearing.157 - Treat repeated "Taken together," as a generator-voice smell. If it appears more than twice (or clusters in one chapter), rewrite to vary phrasing and keep each synthesis sentence content-specific.158 - Vary synthesis openings: "In summary," "Across these studies," "The pattern that emerges," "A key insight," "Collectively," "The evidence suggests," or directly state the conclusion without a synthesis marker.159 - Each synthesis opening should be content-specific, not a template label.160161Rewrite recipe for subsection openers (paper voice, no new facts):162- Delete: `This subsection surveys/argues...` / `In this subsection, we...`163- Replace with a compact opener that does 2–3 of these (no labels; vary across subsections):164 - **Content claim**: the subsection-specific tension/trade-off (optionally with 1–2 embedded citations)165 - **Why it matters**: link the claim to evaluation/engineering constraints (benchmark/protocol/cost/tool access)166 - **Preview**: what you will contrast next and on what lens (A vs B; then evaluation anchors; then limitations)167- Example skeletons (paraphrase; don’t reuse verbatim):168 - Tension-first: `A central tension is ...; ...; we contrast ...`169 - Decision-first: `For builders, the crux is ...; ...`170 - Lens-first: `Seen through the lens of ..., ...`171172### Pass 2 — Terminology normalization173174Role split:175- **Taxonomist**: chooses canonical terms and synonym policy.176- **Integrator**: applies consistent replacements across the draft.177178Targets:179- One concept = one name across sections.180- Headings, tables, and prose use the same canonical terms.181182### Pass 3 — Redundancy pruning (global repetition)183184Role split:185- **Compressor**: collapses repeated boilerplate.186- **Narrative keeper**: ensures removing repetition does not break the argument chain.187188Targets:189- Cross-section repeated intros/outros are removed.190- Only subsection-specific content remains inside subsections.191192## Script193194### Quick Start195196- `uv run python .codex/skills/draft-polisher/scripts/run.py --help`197- `uv run python .codex/skills/draft-polisher/scripts/run.py --workspace <workspace>`198199### All Options200201- `--workspace <dir>`: workspace root202- `--unit-id <U###>`: unit id (optional; for logs)203- `--inputs <semicolon-separated>`: override inputs (rare; prefer defaults)204- `--outputs <semicolon-separated>`: override outputs (rare; prefer defaults)205- `--checkpoint <C#>`: checkpoint id (optional; for logs)206207### Examples208209- First polish pass (creates anchoring baseline `output/citation_anchors.prepolish.jsonl`):210 - `uv run python .codex/skills/draft-polisher/scripts/run.py --workspace <workspace>`211212- Reset the anchoring baseline (only if you intentionally accept citation drift):213 - Delete `output/citation_anchors.prepolish.jsonl`, then rerun the polisher.214215## Acceptance checklist216217- [ ] No `TODO/TBD/FIXME/(placeholder)`.218- [ ] No `…` or `...` truncation.219- [ ] No repeated boilerplate sentence across many subsections.220- [ ] Citation anchoring passes (no cross-subsection drift).221- [ ] Each H3 has at least one cross-paper synthesis paragraph (>=2 citations).222223## Troubleshooting224225### Issue: polishing causes citation drift across subsections226227**Fix**:228- Keep citations inside the same `###` subsection; if restructuring is intentional, delete `output/citation_anchors.prepolish.jsonl` and regenerate a new baseline.229230### Issue: draft polishing is requested before writing approval231232**Fix**:233- Record the relevant approval in `DECISIONS.md` (typically `Approve C2`) before doing prose-level edits.