Paper Writer
Overview
This skill turns a researcher's ideas and materials into submittable paper
prose. One paragraph works; a full manuscript works.
It is not an outline generator and not a writing tutor. It writes the text,
under one non-negotiable condition: every written word must be traceable to
evidence. The author owns the substance (ideas, designs, results,
conclusions); this skill owns the linguistic realization of that substance,
and nothing more.
Hard rules
These hold for every task this skill performs:
- Every factual claim has one of three origins: the user's materials, this
session's verified retrieval, or field common knowledge that carries no
numbers, names, or comparisons. Model memory is never a source.
- Delivered prose contains zero bracketed placeholder tags of any kind. A
claim without a source is resolved by searching, rewriting, or deleting,
never by tagging.
- Concrete details the user did not provide are not written: no invented
scenarios, mechanisms, magnitudes, procedures, or identifiers.
- Claim strength never exceeds evidence strength.
- Real results and planned or expected results are phrased differently and
never mixed.
- Internal planning never leaks into the output.
The full discipline, including the evidence hierarchy and the Evidence Map:
see references/evidence-discipline.md.
When to use this skill
- "Write this idea into an English introduction paragraph."
- "Draft the Discussion in the style of a specific venue."
- "Write a full first draft of the paper."
- "Turn these notes into a Methods section."
- "Produce an abstract."
When NOT to use this skill
- The user wants an Introduction specifically:
intro-drafter owns that
section (this skill routes there automatically).
- The paper is a benchmark paper still being planned:
benchmark-paper-template
first, then return here for the prose.
- Existing prose needs language polishing:
paper-polish.
- The logic skeleton is not settled:
tech-paper-template first.
- The idea itself is unvetted:
idea-evaluator.
Capability check (run once per session)
This skill uses two environment capabilities and degrades honestly when they
are missing:
- Literature search: a scholarly search tool if the harness provides
one; otherwise web search against scholarly indexes; otherwise shell access
to public APIs (Crossref, Semantic Scholar, arXiv, DBLP). If none exist,
operate closed-book: cite only user-supplied references, keep citation
claims at citation level, and disclose in the delivery note that no
independent retrieval was possible.
- Fresh-context sub-agents: used for independent citation verification.
Without them, fall to the same-context grounding rung. The ladder and its
disclosure rules: see references/verification-ladder.md.
The evidence rules above never degrade; only the verification mechanism does,
and any degradation is disclosed.
Workflow
Six phases. A full paper walks all of them explicitly; a single section runs
the same phases in lighter form, but never skips Evidence or Review.
Phase 1: Scope
Settle three things quickly, without interrogating the user:
- Granularity: full paper (produce working files: Evidence Map, chapter
blueprint, then the draft) versus single section or paragraph (inventory
mentally, deliver prose directly).
- Paradigm, judged by research method, not by discipline name:
experiments, benchmarks, models, algorithms mean STEM; textual analysis,
archives, conceptual argument mean humanities; surveys, interviews,
regression, fieldwork mean empirical social science; instrumental
variables, difference-in-differences, panel data mean economics; statutes,
cases, doctrine mean law; a synthesis of existing studies means review.
Unclear: ask one question about the core method and target venue.
- Mode: Draft (default; structural placeholders allowed in tables only)
versus Final (the user said submission-ready; nothing pending is
tolerated).
Phase 2: Evidence
Before any words, lay out what the evidence base contains.
- Inventory the user's materials. When they live in workspace files, read
the files and record paths; do not ask the user to paste what you can
read.
- Build the literature pool with two or three retrieval rounds under
different keywords (target scale: on the order of twenty works for a full
paper; coverage matters, the number does not).
- Identify evidence gaps: which planned claims currently have no L1-L3
source?
Full papers get a written Evidence Map; sections get the same check
mentally. Template, levels, and gate: references/evidence-discipline.md.
Gate: a claim with no L0-L3 source is not written as fact. Search first;
if two or three keyword variants find nothing, rewrite the sentence to drop
the claim, or delete it.
Phase 3: Blueprint
Route by paradigm and section:
| Case |
Route |
| STEM Introduction |
hand off to intro-drafter (six-paragraph internal scaffold) |
| STEM other sections |
references/section-guidance.md content contracts |
| Non-STEM, any section |
plan each paragraph on a CER skeleton: Claim (what should the reader believe), Evidence (what supports it, at which level), Reasoning (why that evidence supports it), Role (motivate, situate, propose, execute, present, interpret, qualify, or connect) |
| Benchmark papers |
plan in benchmark-paper-template, then return here |
For a full paper, produce a chapter blueprint as a working file: section,
role, main judgment, evidence IDs, open gaps. For STEM work, check the logic
chain end to end: limitations feed the key idea, the key idea raises the
challenges, modules answer challenges one to one, contributions cover the
modules. Broken links get fixed in the plan, never papered over in prose.
All of this is internal. None of it appears in the deliverable.
Phase 4: Draft
Write flowing prose from the blueprint, with provenance thinking: before each
paragraph, settle internally what it will claim and which source backs each
claim; a claim with no source is excluded during planning. Write clean text
with no embedded markers of any kind.
Hard writing rules:
- Never generate content from model memory. Three legitimate origins
only (user, retrieval, L0). Uncertain facts get verified through
retrieval; verification failure means rewrite or delete, never tag.
- Evidence level caps claim strength. L1 supports anything; L2 supports
directional summaries; L3 supports citation-level statements only; L4
supports nothing.
- No invented specifics. The five red-flag families (scenarios,
mechanisms, magnitudes, procedures, identifiers) are the ban list;
omission beats plausible invention every time.
- Separate real from planned results. Confirmed results: "we observe",
"results show". Expected or unconfirmed: "the authors report",
"preliminary results suggest".
- Never oversell for the author. Without evidence: "may", "shows
promise for", "is expected to".
- Each section does its own job, once. Introduction makes the reader
care, shows the gap, states the goal. Methods let another researcher
reproduce. Results report observations without why. Discussion explains
why, connects to prior work, admits limits. Conclusion answers the
Introduction. Abstract is a self-contained miniature written last. The
same finding appears in at most three sections.
- No internal scaffolding in the output. CER skeletons, paradigm calls,
chain checks all stay silent.
Phase 4.5: Red flags (stop signals while writing)
| The thought |
The correct move |
| "I remember this scale has 10 items" |
Stop. Uncertain. Omit the item count |
| "I recall this paper used a meta-analysis" |
Stop. No full text seen. Citation level only: "X et al. (Year) studied Y" |
| "I remember this case's citation number" |
Stop. One digit off is wrong. Omit the identifier |
| "This material's conductivity is roughly..." |
Stop. No cross-material comparisons unless the user gave the numbers |
| "This country's policy works like..." |
Stop. Institutional detail only as the user described it |
| "Background should mention the field's development" |
Stop. Background claims beyond L0 need L1-L3 sources. Search; nothing found means do not write it |
| "Readers expect consent details here" |
Stop. Not provided means omitted |
| "A complete-looking table is more professional" |
Stop. -- in a cell is 100 times safer than an invented value |
| "This Discussion should be fuller" |
Stop. Enrich with evidence, not invented mechanisms. Short and accurate beats long and hollow |
| "This method could apply to solar farms / autonomous driving..." |
Stop. The user named no such scenario |
| "It works because of spectral decomposition / gradient coupling..." |
Stop. The user described no such mechanism |
| "Targets differ tenfold / fewer than ten pixels / hundreds of classes" |
Stop. The user gave no such magnitudes |
| "A concrete failure case would illustrate this" |
Stop. No invented domain vignettes. Use the user's own description |
Phase 5: Review
Three parts, in order:
- Source recall check plus inline checklist. Walk
references/verification-checklist.md: facts, evidence matching,
cross-section discipline, AI-trace scan, formulas and tables, delivery
format.
- Independent citation verification. Mandatory for full papers, Final
mode, or three or more citations; a fresh-context sub-agent receives only
the prose and the reference list and verifies every entry through
retrieval. Statuses, resolution routes, and the degradation ladder:
references/verification-ladder.md. Delivery waits until no unresolved
problem entries remain.
- Know what review cannot do. Same-context review reliably catches
mechanical issues (numbering, format, cross-section duplication). It
cannot reliably catch its own semantic fabrication; a model that invented
a scenario will confirm that scenario on re-reading. The real defense is
the write-time evidence gate in Phases 2-4; the review phase is the
backstop, and the citation layer gets the independent pass precisely
because it is the one layer that can be mechanically outsourced.
Phase 6: Deliver
Deliverable and prohibitions: references/prose-delivery.md.
- Single section or paragraph: prose in the conversation, References list
when citations are present, at most three lines of notes.
- Full paper: the manuscript written to a workspace file (Markdown or
LaTeX), References included; Evidence Map and blueprint remain separate
working files, available on request, never embedded in the manuscript.
- Any capability degradation (no retrieval, no sub-agents) is stated in the
note.
Proactive literature searching
Do not wait for the user to hand over every reference. Sparse citation reads
as an opinion piece and reviewers reject it.
Reference density (guidance, not quota): a full Introduction typically weaves
15-25 references; a single gap or background paragraph 3-6; Related Work
15-30; Methods 3-8 (cited methods, datasets, protocols); Discussion 5-15; a
full paper on the order of 20-40.
Search before writing, not after. Before drafting any section, run three
to five retrieval rounds under different angles: core terminology; method or
model names; application domain; venue or author names. The retrieved works
are the material the argument is organized around, not decoration attached
afterwards.
Retrieval returns metadata (L3), which supports citation-level statements
only: "Recent work by Author et al. addressed X using Y" is fine; "their
method relies on a three-stage pipeline" or a specific number is not, unless
the user supplied the abstract or full text. Need full-text-level content
with only L3 in hand: write the citation-level version and tell the user in
the note which source would unlock the stronger sentence.
If a drafted section comes out visibly under-cited, go back and search
another round, then weave real findings in. Never bridge the gap with a
placeholder.
Citations and references
Default style is numeric citation-sequence with a References section that
matches the text bidirectionally. Entry formats, author rules, missing-field
handling, and the alternate styles (APA, Chicago, Harvard, Vancouver):
references/citation-style.md.
Output language
Follow an explicit language request first. If the target is a Chinese
journal or a Chinese thesis, write Chinese academic prose; otherwise default
to English. Never infer the output language from the language the user typed
in: describing ideas in Chinese and submitting in English is the normal
case.
Boundaries with sibling skills
intro-drafter: Introduction prose (this skill routes STEM Introductions
there).
paper-polish: polishing existing prose, including Chinese-to-English.
tech-paper-template: the pre-writing logic skeleton.
benchmark-paper-template: planning benchmark papers.
pre-submission-reviewer: reviewer-style audit of the finished draft.
idea-evaluator: whether the idea is worth pursuing at all.
If drafting reveals that the problem is not "how to write it" but "does the
core claim hold", say so plainly in the delivery note and point to
pre-submission-reviewer or idea-evaluator. Never write a defense for a
claim the user's own data undermines.
1---2name: paper-writer3description: Drafts publishable paper prose from the author's own materials, from a single paragraph to a full manuscript, across STEM and non-STEM fields. Every factual claim traces to user input, verified retrieval, or field common knowledge; citations pass an independent verification ladder; delivery is clean prose with zero placeholder tags. Use when the user asks to write a section, turn an idea or results into paper text, or draft a full paper.4license: CC-BY-NC-SA-4.05---6
7# Paper Writer
8
9## Overview
10
11This skill turns a researcher's ideas and materials into submittable paper
12prose. One paragraph works; a full manuscript works.
13
14It is not an outline generator and not a writing tutor. It writes the text,
15under one non-negotiable condition: every written word must be traceable to
16evidence. The author owns the substance (ideas, designs, results,
17conclusions); this skill owns the linguistic realization of that substance,
18and nothing more.
19
20## Hard rules
21
22These hold for every task this skill performs:
23
241. Every factual claim has one of three origins: the user's materials, this
25 session's verified retrieval, or field common knowledge that carries no
26 numbers, names, or comparisons. Model memory is never a source.
272. Delivered prose contains zero bracketed placeholder tags of any kind. A
28 claim without a source is resolved by searching, rewriting, or deleting,
29 never by tagging.
303. Concrete details the user did not provide are not written: no invented
31 scenarios, mechanisms, magnitudes, procedures, or identifiers.
324. Claim strength never exceeds evidence strength.
335. Real results and planned or expected results are phrased differently and
34 never mixed.
356. Internal planning never leaks into the output.
36
37The full discipline, including the evidence hierarchy and the Evidence Map:
38see references/evidence-discipline.md.
39
40## When to use this skill
41
42- "Write this idea into an English introduction paragraph."
43- "Draft the Discussion in the style of a specific venue."
44- "Write a full first draft of the paper."
45- "Turn these notes into a Methods section."
46- "Produce an abstract."
47
48## When NOT to use this skill
49
50- The user wants an Introduction specifically: `intro-drafter` owns that
51 section (this skill routes there automatically).
52- The paper is a benchmark paper still being planned: `benchmark-paper-template`
53 first, then return here for the prose.
54- Existing prose needs language polishing: `paper-polish`.
55- The logic skeleton is not settled: `tech-paper-template` first.
56- The idea itself is unvetted: `idea-evaluator`.
57
58## Capability check (run once per session)
59
60This skill uses two environment capabilities and degrades honestly when they
61are missing:
62
63- **Literature search**: a scholarly search tool if the harness provides
64 one; otherwise web search against scholarly indexes; otherwise shell access
65 to public APIs (Crossref, Semantic Scholar, arXiv, DBLP). If none exist,
66 operate closed-book: cite only user-supplied references, keep citation
67 claims at citation level, and disclose in the delivery note that no
68 independent retrieval was possible.
69- **Fresh-context sub-agents**: used for independent citation verification.
70 Without them, fall to the same-context grounding rung. The ladder and its
71 disclosure rules: see references/verification-ladder.md.
72
73The evidence rules above never degrade; only the verification mechanism does,
74and any degradation is disclosed.
75
76## Workflow
77
78Six phases. A full paper walks all of them explicitly; a single section runs
79the same phases in lighter form, but never skips Evidence or Review.
80
81### Phase 1: Scope
82
83Settle three things quickly, without interrogating the user:
84
85- **Granularity**: full paper (produce working files: Evidence Map, chapter
86 blueprint, then the draft) versus single section or paragraph (inventory
87 mentally, deliver prose directly).
88- **Paradigm**, judged by research method, not by discipline name:
89 experiments, benchmarks, models, algorithms mean STEM; textual analysis,
90 archives, conceptual argument mean humanities; surveys, interviews,
91 regression, fieldwork mean empirical social science; instrumental
92 variables, difference-in-differences, panel data mean economics; statutes,
93 cases, doctrine mean law; a synthesis of existing studies means review.
94 Unclear: ask one question about the core method and target venue.
95- **Mode**: Draft (default; structural placeholders allowed in tables only)
96 versus Final (the user said submission-ready; nothing pending is
97 tolerated).
98
99### Phase 2: Evidence
100
101Before any words, lay out what the evidence base contains.
102
1031. Inventory the user's materials. When they live in workspace files, read
104 the files and record paths; do not ask the user to paste what you can
105 read.
1062. Build the literature pool with two or three retrieval rounds under
107 different keywords (target scale: on the order of twenty works for a full
108 paper; coverage matters, the number does not).
1093. Identify evidence gaps: which planned claims currently have no L1-L3
110 source?
111
112Full papers get a written Evidence Map; sections get the same check
113mentally. Template, levels, and gate: references/evidence-discipline.md.
114
115**Gate**: a claim with no L0-L3 source is not written as fact. Search first;
116if two or three keyword variants find nothing, rewrite the sentence to drop
117the claim, or delete it.
118
119### Phase 3: Blueprint
120
121Route by paradigm and section:
122
123| Case | Route |
124|---|---|
125| STEM Introduction | hand off to `intro-drafter` (six-paragraph internal scaffold) |
126| STEM other sections | references/section-guidance.md content contracts |
127| Non-STEM, any section | plan each paragraph on a CER skeleton: Claim (what should the reader believe), Evidence (what supports it, at which level), Reasoning (why that evidence supports it), Role (motivate, situate, propose, execute, present, interpret, qualify, or connect) |
128| Benchmark papers | plan in `benchmark-paper-template`, then return here |
129
130For a full paper, produce a chapter blueprint as a working file: section,
131role, main judgment, evidence IDs, open gaps. For STEM work, check the logic
132chain end to end: limitations feed the key idea, the key idea raises the
133challenges, modules answer challenges one to one, contributions cover the
134modules. Broken links get fixed in the plan, never papered over in prose.
135
136All of this is internal. None of it appears in the deliverable.
137
138### Phase 4: Draft
139
140Write flowing prose from the blueprint, with provenance thinking: before each
141paragraph, settle internally what it will claim and which source backs each
142claim; a claim with no source is excluded during planning. Write clean text
143with no embedded markers of any kind.
144
145Hard writing rules:
146
1471. **Never generate content from model memory.** Three legitimate origins
148 only (user, retrieval, L0). Uncertain facts get verified through
149 retrieval; verification failure means rewrite or delete, never tag.
1502. **Evidence level caps claim strength.** L1 supports anything; L2 supports
151 directional summaries; L3 supports citation-level statements only; L4
152 supports nothing.
1533. **No invented specifics.** The five red-flag families (scenarios,
154 mechanisms, magnitudes, procedures, identifiers) are the ban list;
155 omission beats plausible invention every time.
1564. **Separate real from planned results.** Confirmed results: "we observe",
157 "results show". Expected or unconfirmed: "the authors report",
158 "preliminary results suggest".
1595. **Never oversell for the author.** Without evidence: "may", "shows
160 promise for", "is expected to".
1616. **Each section does its own job, once.** Introduction makes the reader
162 care, shows the gap, states the goal. Methods let another researcher
163 reproduce. Results report observations without why. Discussion explains
164 why, connects to prior work, admits limits. Conclusion answers the
165 Introduction. Abstract is a self-contained miniature written last. The
166 same finding appears in at most three sections.
1677. **No internal scaffolding in the output.** CER skeletons, paradigm calls,
168 chain checks all stay silent.
169
170### Phase 4.5: Red flags (stop signals while writing)
171
172| The thought | The correct move |
173|---|---|
174| "I remember this scale has 10 items" | Stop. Uncertain. Omit the item count |
175| "I recall this paper used a meta-analysis" | Stop. No full text seen. Citation level only: "X et al. (Year) studied Y" |
176| "I remember this case's citation number" | Stop. One digit off is wrong. Omit the identifier |
177| "This material's conductivity is roughly..." | Stop. No cross-material comparisons unless the user gave the numbers |
178| "This country's policy works like..." | Stop. Institutional detail only as the user described it |
179| "Background should mention the field's development" | Stop. Background claims beyond L0 need L1-L3 sources. Search; nothing found means do not write it |
180| "Readers expect consent details here" | Stop. Not provided means omitted |
181| "A complete-looking table is more professional" | Stop. `--` in a cell is 100 times safer than an invented value |
182| "This Discussion should be fuller" | Stop. Enrich with evidence, not invented mechanisms. Short and accurate beats long and hollow |
183| "This method could apply to solar farms / autonomous driving..." | Stop. The user named no such scenario |
184| "It works because of spectral decomposition / gradient coupling..." | Stop. The user described no such mechanism |
185| "Targets differ tenfold / fewer than ten pixels / hundreds of classes" | Stop. The user gave no such magnitudes |
186| "A concrete failure case would illustrate this" | Stop. No invented domain vignettes. Use the user's own description |
187
188### Phase 5: Review
189
190Three parts, in order:
191
1921. **Source recall check plus inline checklist.** Walk
193 references/verification-checklist.md: facts, evidence matching,
194 cross-section discipline, AI-trace scan, formulas and tables, delivery
195 format.
1962. **Independent citation verification.** Mandatory for full papers, Final
197 mode, or three or more citations; a fresh-context sub-agent receives only
198 the prose and the reference list and verifies every entry through
199 retrieval. Statuses, resolution routes, and the degradation ladder:
200 references/verification-ladder.md. Delivery waits until no unresolved
201 problem entries remain.
2023. **Know what review cannot do.** Same-context review reliably catches
203 mechanical issues (numbering, format, cross-section duplication). It
204 cannot reliably catch its own semantic fabrication; a model that invented
205 a scenario will confirm that scenario on re-reading. The real defense is
206 the write-time evidence gate in Phases 2-4; the review phase is the
207 backstop, and the citation layer gets the independent pass precisely
208 because it is the one layer that can be mechanically outsourced.
209
210### Phase 6: Deliver
211
212Deliverable and prohibitions: references/prose-delivery.md.
213
214- Single section or paragraph: prose in the conversation, References list
215 when citations are present, at most three lines of notes.
216- Full paper: the manuscript written to a workspace file (Markdown or
217 LaTeX), References included; Evidence Map and blueprint remain separate
218 working files, available on request, never embedded in the manuscript.
219- Any capability degradation (no retrieval, no sub-agents) is stated in the
220 note.
221
222## Proactive literature searching
223
224Do not wait for the user to hand over every reference. Sparse citation reads
225as an opinion piece and reviewers reject it.
226
227Reference density (guidance, not quota): a full Introduction typically weaves
22815-25 references; a single gap or background paragraph 3-6; Related Work
22915-30; Methods 3-8 (cited methods, datasets, protocols); Discussion 5-15; a
230full paper on the order of 20-40.
231
232**Search before writing, not after.** Before drafting any section, run three
233to five retrieval rounds under different angles: core terminology; method or
234model names; application domain; venue or author names. The retrieved works
235are the material the argument is organized around, not decoration attached
236afterwards.
237
238Retrieval returns metadata (L3), which supports citation-level statements
239only: "Recent work by Author et al. addressed X using Y" is fine; "their
240method relies on a three-stage pipeline" or a specific number is not, unless
241the user supplied the abstract or full text. Need full-text-level content
242with only L3 in hand: write the citation-level version and tell the user in
243the note which source would unlock the stronger sentence.
244
245If a drafted section comes out visibly under-cited, go back and search
246another round, then weave real findings in. Never bridge the gap with a
247placeholder.
248
249## Citations and references
250
251Default style is numeric citation-sequence with a References section that
252matches the text bidirectionally. Entry formats, author rules, missing-field
253handling, and the alternate styles (APA, Chicago, Harvard, Vancouver):
254references/citation-style.md.
255
256## Output language
257
258Follow an explicit language request first. If the target is a Chinese
259journal or a Chinese thesis, write Chinese academic prose; otherwise default
260to English. Never infer the output language from the language the user typed
261in: describing ideas in Chinese and submitting in English is the normal
262case.
263
264## Boundaries with sibling skills
265
266- `intro-drafter`: Introduction prose (this skill routes STEM Introductions
267 there).
268- `paper-polish`: polishing existing prose, including Chinese-to-English.
269- `tech-paper-template`: the pre-writing logic skeleton.
270- `benchmark-paper-template`: planning benchmark papers.
271- `pre-submission-reviewer`: reviewer-style audit of the finished draft.
272- `idea-evaluator`: whether the idea is worth pursuing at all.
273
274If drafting reveals that the problem is not "how to write it" but "does the
275core claim hold", say so plainly in the delivery note and point to
276`pre-submission-reviewer` or `idea-evaluator`. Never write a defense for a
277claim the user's own data undermines.