Paper Deep Dive
Short description: PDF to deep paper notes workflow. Historical source: Feishu wiki AI Research Skills; current durable delivery is Notion-first.
Use this skill when the user wants a complete paper reading workflow from PDF to stable notes, paper card, assets, and shareable summaries.
Single Source of Truth
This skill is the only canonical delivery standard for single-paper deep dives. Other skills may trigger, route, write to Feishu/Notion/Obsidian, preserve Chinese wording, or distill the finished deep dive into a wiki, but they must not define a second deep-dive structure or mark a package complete under weaker rules.
When any user request says deep dive, 深读, 详细解析, detailed-read, dive into, 精读这篇论文, or requests a complete 原文中译稿, treat it as this full workflow unless the user explicitly asks for a lighter artifact such as quick summary, paper card only, partial translation, or English manuscript only. A lighter artifact must be labeled as such and must not be called a compliant deep dive.
The completion standard is product-level, not effort-level. A good summary or a partial translation is still incomplete. A package is compliant when the main page contains the complete Chinese manuscript and its required 精读部分, and figures/captions, formulas, references, hierarchy, and target-platform read-back verification all pass the delivery gate below. Child pages are not part of the default completion gate.
Canonical Platform Delivery Gate
Also use research-doc-workflow and notion-doc-workflow for new deep dives. Notion is the canonical durable target because Feishu storage is limited. Keep one Markdown-first manuscript and close-reading structure; only Notion hierarchy, native equations/captions, uploads, editable trees, citation chrome, and write verification need platform adaptation. Do not create or maintain a parallel Feishu deep-dive copy unless the user explicitly requests a legacy repair or migration.
Canonical Paper Card Gate
When this workflow creates or modifies a paper card, also use paper-card-delivery. That skill is the canonical standard for official-source verification, fixed card format, figure/caption selection, and structural validation. Do not finalize the card from this deep-dive skill alone.
Canonical Chinese Technical Writing Gate
For 原文中译稿, 完整中文译稿, 精读稿, Chinese figure captions, Chinese table captions/notes, main-entry summaries, and Chinese paper-card prose, also use chinese-technical-writing. Preserve official English source text in 英文原文稿, formulas, named architectures/models/methods such as Transformer and DINO, dataset names, symbols, and table cell content. Translate generic technical concepts into Chinese, but do not translate official names or force them into awkward Chinese names.
Borrowed Method Layer
This is the user's canonical deep-dive workflow. It may borrow useful reading methods from downloaded skills, but those skills do not own the final package semantics or target-platform representation.
- From
nature-reader, borrow the source-map-first method: build stable source
block IDs such as S001, F001, T001, preserve original / Chinese
correspondence, keep figure/table captions attached to the relevant text, and
record uncertainty instead of guessing.
- Do not publish the external
nature-reader artifact contract as-is. A lone paper.md, source_map.json, translation_notes.md, or bilingual reader does not replace the required self-contained main page. These files may be intermediate material or part of an explicitly requested Obsidian package.
- Use the source map as an internal scaffold for the main-page
原文中译稿 and source-grounded 精读部分; use it for an 英文原文稿 only when the user explicitly requests that optional artifact. The final deliverable must use the selected platform's native images/captions, formulas, hierarchy, and read-back verification through research-doc-workflow.
- In the main-page
精读部分, short bilingual source snippets or block IDs may be included when they clarify a key claim, equation, or figure, but this section remains analytical material rather than a second translation.
Mechanism Interrogation Gate
Treat deep reading as reconstruction followed by audit. First recover the strongest version of the authors' logic; only then test whether the problem is important, the assumptions are defensible, the design addresses the stated bottleneck, the evidence excludes relevant alternatives, and the conclusion stays within the evidence boundary. Do not confuse skepticism with automatic rejection.
Before calling the analytical close reading complete, establish all of the following from the source package:
- the concrete prior-method bottleneck and its proposed causal explanation;
- the key assumptions and the falsifiable predictions they imply;
- a design-to-mechanism chain:
design -> changed information/constraint/optimization -> predicted effect;
- the decisive experiment, control, or ablation for each central mechanism claim;
- plausible alternative explanations that the experiments do or do not rule out;
- counterfactual predictions for removing, replacing, or simplifying a module;
- the data, scene, supervision, or optimization conditions under which the method should fail;
- the smallest credible alternative design and the next question that would discriminate between explanations.
Keep 作者声称, 实验支持, 我们的推断, and 尚未验证 distinguishable in analytical notes. These are evidence statuses, not mandatory repeated headings. Do not inject this analysis into the source-faithful 原文中译稿; it belongs in the main-page 精读部分, including the editable tree and analytical close reading.
When To Use
- Reading a new paper deeply rather than only summarizing it.
- Converting the verified paper source into one self-contained main page: complete faithful Chinese translation plus a clearly separated 精读部分 containing the paper card, editable paper-analysis tree, source-order close reading, and mechanism synthesis.
- Preparing Notion paper pages/notes, or repairing an existing legacy page in another platform when explicitly requested.
- Auditing whether figures, claims, assets, and citations are complete.
If the user says deep dive, 深读, 详细解析, dive into, or asks to deeply read a single paper, use this full workflow by default. Do not downgrade it to a quick summary, paper card only, or close-reading note only unless the user explicitly asks for a lighter output.
Non-Negotiable Deliverable
For papers with an accessible official PDF or full-paper HTML, the sole mandatory durable artifact is the Notion main page. Its primary body is the complete faithful 原文中译稿 in source order, followed by a clearly separated, embedded section titled 精读稿. That section contains the Paper Card, editable 论文解析树, source-order analytical close reading, and integrated mechanism synthesis.
<paper short name>|英文原文稿 and a standalone <paper short name>|精读稿 child page are optional artifacts. Do not create either by default. Create them only when the user explicitly requests them or supplies an existing hierarchy that must be preserved during migration/repair. Their absence is never a deep-dive completion failure.
If the PDF can be downloaded or viewed, assume the complete Chinese manuscript can be produced by MinerU extraction plus official HTML / LaTeX / PDF verification. Do not use context length, page length, one-turn time, target-page size, translation workload, or "current tool path" as reasons to downgrade it into a section summary, structured outline, selected excerpts, or partial translation. Chunk the paper by sections, append incrementally, and continue until the main page is complete.
Only three conditions justify not producing the complete main-page Chinese manuscript: the full paper source is inaccessible, reproduction is blocked by a clear licensing/copyright constraint, or the user explicitly asks for a lighter / partial artifact. Otherwise, an incomplete 原文中译稿 is work in progress, not a compliant deep dive.
Source Acquisition Priority
Before running MinerU or writing any target page/note, build the complete source package, including every official or source-derived HTML version that can be found. Search for arXiv HTML (https://arxiv.org/html/<id> and versioned variants), publisher/proceedings HTML, OpenReview/forum HTML, CVF/open-access HTML, and project-page paper HTML when present. HTML is not optional source decoration: it is often the best source for section hierarchy, MathML / TeX annotations, figure/table nodes, captions, references, and supplementary links.
For papers that have an arXiv version, search and inspect the arXiv landing page, arXiv PDF, arXiv HTML, and LaTeX source before using a conference or publisher PDF as the main extraction source. Prefer arXiv HTML / LaTeX for structure, formulas, captions, references, and appendix discovery; prefer arXiv PDF for page layout, figure placement, and visual cross-checking. The reason is practical: many official conference PDFs omit appendices or supplementary sections, while arXiv often preserves the fuller manuscript.
Use venue/publisher pages and their official article/proceedings HTML for official metadata, acceptance venue, project/code links, supplement links, and cross-checking, but do not assume the venue PDF is the complete manuscript. If the only PDF initially found is from CVF, ACM, IEEE, Springer, PMLR, OpenReview, NeurIPS, a conference proceedings site, or a publisher landing page, explicitly search for separate supplementary, appendix, supp, supplemental material, additional material, SM, or PDF supplementary files on the same page, the official HTML page, the proceedings page, OpenReview, the project page, author/lab page, and arXiv.
When HTML and PDF disagree, treat official HTML / LaTeX as the first authority for source text, section order, formulas, captions, and references when it clearly preserves the paper source; cross-check figure placement, page layout, missing appendices, and visual assets against the PDF and supplement. Record meaningful discrepancies instead of silently choosing one source.
If a separate appendix/supplementary PDF exists, it is part of the deep-dive source package unless the user explicitly excludes it. Download it alongside the main PDF, include its sections, figures, tables, formulas, captions, algorithms, and references in the source map, and reflect it in the main-page 原文中译稿. If an optional English manuscript is explicitly requested, include the same source coverage there. A deep dive based only on a conference main PDF is incomplete when a separate supplement exists and has not been inspected.
Supplementary / appendix PDF extraction priority (mandatory): when a separate supplementary or appendix PDF has no usable official HTML (typical for project-page or ACM/IEEE “Supplementary Material” PDFs), the required first conversion path is MinerU via .tools/mineru-md.sh (same Local MinerU Extraction rules as the main PDF). Do not use PyMuPDF / pdfplumber / raw get_text / page screenshots as the primary manuscript scaffold for that supplement. After MinerU, repair formulas/captions/tables against the PDF (and against HTML/LaTeX only if a supplement HTML/LaTeX source actually exists). PyMuPDF and similar tools are allowed only as a fallback after MinerU fails, or for narrow visual checks (page count, figure crop verification), not as the default text path for supplements.
Record the source package in task notes and compact source metadata: arXiv ID/version when available, arXiv PDF status, arXiv HTML URL/status, LaTeX source status, publisher/proceedings/OpenReview/CVF HTML URL/status, venue/publisher PDF status, supplementary/appendix PDF status, project/code links, and extraction date. If no HTML version is found, state the searched routes briefly; do not write HTML not checked. If no supplement is found, state the searched routes briefly; do not write supplement not checked.
Workflow
Imported PDF Translation Repair Path
When the user provides an arXiv PDF or paper title and asks to download, run
pdf2zh, and import the Chinese PDF into Notion, also use
paper-pdf2zh-notion-import. That
skill owns the direct Notion desktop import, title-based PDF renaming, local
artifact cleanup, and the mandatory two-round
repair/read-back cycle. The imported PDF is always draft material; a successful
Notion upload is not a completed 原文中译稿.
When the user has already translated a paper PDF with pdf2zh-next and manually imported the resulting PDF into Notion or Feishu, treat the imported page as a draft translation, not as a completed 原文中译稿. This path is an accepted entry point and does not require re-running PDF translation, but it must pass the same source-fidelity and platform verification gates before delivery.
One-line workflow trigger: 帮我 pdf2zh 并导入 Notion is sufficient to invoke the full import workflow. Do not ask the user to restate the repair rules. Resolve the latest arXiv PDF, run the local conversion, import the Chinese PDF directly into the supplied Notion parent, then perform the two-round repair/read-back cycle below.
Imported-page delivery rule: The existing imported Chinese page is the primary deliverable, matching the standard deep-dive structure. Do not require creation of 英文原文稿 or standalone 精读稿 child pages unless the user explicitly asks for them. Verify the Chinese manuscript against the latest official PDF and available HTML/LaTeX, repair source-order structure, formulas, figures and native captions, tables, appendices, references, and body citations, add or repair the required main-page 精读部分, then perform target-platform read-back verification.
- A. Identify the source. Locate the imported page and source PDF. Preserve imported figures, tables, and page order as draft material, then compare them against the source PDF and, when available, official HTML / LaTeX.
- B. Fix identity and navigation. Rename the page to the verified Chinese paper title only. Keep the official English title in the opening block. Add the latest arXiv PDF, Project Page, and Code links above the abstract when available. Do not add
英文原文稿 or standalone 精读稿 child links unless explicitly requested; never infer URLs from a filename.
- C. Repair structure. Restore heading levels from the paper's numbered hierarchy, reconnect PDF-split paragraphs, remove duplicate headers/footers and conversion artifacts, and keep figures/tables near their source positions.
- C1. Normalize numbered headings at every depth. Use the numeric prefix to determine hierarchy, not the importer’s visual level: a one-part prefix such as
N. (3., 4.) is a top-level section; a two-part prefix such as N.M. (3.1., 4.2.) is its subsection; a three-part prefix such as N.M.K. (3.1.1., 4.2.1.) is its sub-subsection; continue the same rule for deeper prefixes. Normalize full-width heading punctuation . (U+FF0E) to the ASCII period . before parsing, so 4.2.方法 becomes 4.2. 方法. In ordinary structural text, normalize full-width slash / (U+FF0F) to /, full-width hyphen-minus - (U+FF0D) to -, and full-width brackets [] (U+FF3B/U+FF3D) to [] when they are citation, list, or link delimiters. Protect LaTeX, code, URLs, file paths, and existing backslash-escaped sequences before this cleanup, then restore them exactly; never perform a blind global replacement. Do not replace unrelated Chinese punctuation in ordinary prose. Default to ## 3. Section title, ### 3.1. Subsection title, and #### 3.1.1. Sub-subsection title. If the source consistently omits punctuation (3 Title, 3.1 Subtitle), preserve that style for the whole manuscript; never mix punctuated and unpunctuated forms or place a deeper numeric prefix above its parent.
- D. Repair formulas. Restore inline formulas as native inline equations or exact
$...$ LaTeX, restore display equations and numbering, and sample early, middle, formula-heavy, and appendix sections.
- E. Repair figures and tables. Convert figure captions to native caption fields and keep the complete translated source caption without a duplicate paragraph or agent-written summary. Preserve English table cells, translate only table titles/notes, and use an original PDF/HTML screenshot when table layout is unreliable.
- E0. Caption parser check. Before writing a native Notion/Feishu image caption, verify that escaped citation delimiters such as
\[33\] will not be parsed as block syntax. If the caption parser cannot preserve them, convert only those citation markers inside the caption to ordinary visible parentheses while keeping body citations and reference numbering unchanged; then verify that the caption is stored on the image block and no duplicate caption paragraph remains.
- E1. Repair imported table corruption. Inspect every imported table block and its surrounding text for OCR/PDF conversion residue: duplicated Markdown pipe tables after a native table, broken rows or columns, repeated cell fragments, garbled characters, malformed separators, and captions fused to table data. Keep one authoritative editable table, reconstruct rows/cells from the official PDF/HTML when the extracted structure is reliable, and otherwise add an official PDF/HTML table screenshot as the visual authority. Remove duplicate pseudo-tables and keep exactly one translated table title/note attached to the table.
- E2. Author contacts and resource links. Author email addresses must be written as ordinary visible text, not intentionally wrapped in Markdown links,
mailto: links, or code formatting. Notion may auto-link a bare email during rendering; do not add an explicit link or change its visible text to code merely to fight that platform behavior. Only in the opening resource block of a deep-dive Chinese manuscript, omit the heading 来源 and display each link's URL as its link text, for example [https://arxiv.org/pdf/<id>](https://arxiv.org/pdf/<id>). This URL-as-label rule does not apply to English manuscripts, paper cards, or general research documents.
- F. Repair citations. A Chinese-manuscript reference title may be translated, but authors, venue, publisher, year, volume/issue, pages, DOI/arXiv identifiers, URLs, and other bibliographic metadata remain in the source language. Keep
[n] labels, one reference per paragraph, and restore verified PDF URLs. Body citation links must use the same URL map.
- G. Verify completion. Run the full main-page completion gate and record unresolved figure, table, formula, reference, URL, hierarchy, Paper Card, editable-tree, or close-reading issues. A manually imported page is complete only after repair and fetch/read-back verification; the absence of optional child artifacts is not a failure.
Two-round repair is mandatory for imported pages. Round 1 fixes structure and source fidelity against HTML/LaTeX/PDF. Round 2 starts from a fresh Notion read-back and independently audits formulas, inline math, figure/caption placement, table integrity, references, body citation links, appendices, and conversion residue. A first-pass upload or a single visual scan is never a completed delivery.
- Capture source metadata and source package inventory: title, authors, year, venue, DOI/arXiv, arXiv PDF URL/status, arXiv HTML URL/status, LaTeX source status, publisher/proceedings/OpenReview/CVF HTML URL/status, venue/publisher PDF URL, supplementary/appendix PDF URL(s), project/code links, local source paths, and extraction date.
- Extract or parse the complete source package to inspectable Markdown when tooling is available; preserve figure references and equation context. On this machine, use MinerU as the default PDF-to-Markdown path before building deep-dive artifacts, but parse official HTML / LaTeX first when available for source structure, formulas, captions, references, and appendix coverage. Parse the arXiv/full manuscript first when available, then parse any separate supplement/appendix PDF with the same priority: official supplement HTML/LaTeX if present, otherwise MinerU on the supplement PDF before any other PDF text extractor.
- Check the MinerU conversion draft against official HTML whenever HTML exists, especially arXiv HTML for arXiv papers and publisher/proceedings HTML for non-arXiv papers. This check is mandatory, not optional. Repair section order, paragraph continuity, formulas, figures, tables, captions, appendices, body citations, and references before publishing.
- Build a source map inspired by
nature-reader: stable block IDs for body text, figures, tables, captions, equations, appendices, and references; page / section location; extraction confidence; and links between first figure/table mention and the visual asset.
- Build a verified English source scaffold internally from official HTML/LaTeX/PDF. It must be complete enough to support paragraph-level translation and audit, but it is a working artifact, not a required reader-facing page. If the user explicitly requests
<paper short name>|英文原文稿, publish the complete original English text in source order and verify it independently.
- English source correction gate (mandatory before Chinese translation): for arXiv papers, re-open arXiv HTML and correct the internal English scaffold against it. Check section order, paragraph continuity, formulas, figure/table captions, appendix/supplement coverage, body citations, and References. Prefer arXiv HTML for text/structure/formulas and use PDF for layout/visual cross-check. For non-arXiv papers, use the best official HTML; if none exists, correct against LaTeX/PDF and record that HTML correction was impossible. This gate applies even when no English child page is published.
- Create the complete faithful Chinese manuscript directly in the main page from the corrected source scaffold. It must preserve section hierarchy, paragraph correspondence, formulas, figure/table positions, citations, captions, references, appendices/supplements, and layout structure as much as the target editor allows. Translate the paper body, figure captions, table captions/notes (表注), appendix/supplement prose, and explanatory text into Chinese, but keep table cell content in the original English and keep References / bibliography entries source-faithful with the
[n] … . URL [url](url) PDF-link contract. A partial translation is allowed only as a clearly marked WIP state.
- Chinese terminology correction gate (mandatory after the Chinese manuscript draft exists): do a dedicated second pass over
原文中译稿 for terminology only. Verify key method/model/dataset/loss/module terms are consistent; keep named architectures, models, methods, datasets, and official components such as Transformer and DINO in their official English form; add a Chinese gloss only when it improves comprehension; and translate generic technical concepts instead of leaving avoidable English phrase islands. Fix inconsistent renderings of the same term across sections. Do not mark the Chinese manuscript complete until this terminology pass is done.
8b. Reference / citation verification gate (mandatory for the main manuscript): build a single [n] → PDF URL map (prefer arXiv PDF), apply it to References (. URL [url](url)) and body ([[n](url)]), then verify URL correctness and cross-consistency as specified in Reference and Citation Link Contract. If an optional English manuscript is published, apply and verify the same map there. For an imported PDF2ZH Chinese manuscript, the cited paper title may be translated, but all other bibliographic fields must remain source-faithful.
- Create or update the main reader-facing deep-dive page. The primary body is the complete
原文中译稿 and the default reading surface. Its opening must follow a paper-like title block before the abstract: official English title, Chinese title, original English author list and affiliations, then separate verified links for the latest arXiv PDF, Project Page, and Code. Continue with the source-faithful abstract and manuscript in normal paper order. After the full manuscript, add a clearly separated 精读部分; do not interleave analysis with the translation. Do not create or link child artifacts unless explicitly requested.
- Create an editable
论文解析树 that follows the paper's actual reasoning: problem -> concrete bottleneck -> key assumption -> design/mechanism -> changed information or constraint -> predicted effect -> decisive evidence -> boundary. Make the information-flow view (what passes between modules) and the causal-chain view (why the design should change the result) distinguishable. Add losses/training, datasets/evaluation, limitations, and user research implications where they clarify this logic rather than as disconnected inventory branches. Use a native Feishu mind map for Feishu, a structured page/database or supported embedded artifact for Notion, and Mermaid/Canvas plus a searchable linked outline for Obsidian. Do not substitute a static screenshot when an editable representation is available.
- In the main-page
精读部分, create the Paper Card first, then the editable 论文解析树, then a source-order analytical close reading and integrated mechanism synthesis. Follow the paper's own section order and local context: Abstract / Introduction, numbered sections, named subsections, conclusion, then appendices or supplementary material. For each part, explain which claim it advances, why that step is needed, what mechanism or evidence is introduced, and what remains unresolved. Do not insert a repeated per-section heading or paragraph such as "what this means for my world-model research" / "对你的 world model 研究意味着什么". Put user-specific implications and future project ideas only in the final synthesis. This is interpretation, not part of the source-faithful translation.
- Inside
精读部分, use ### or lower-impact paragraph/list structure for source-order close-reading subsections. Do not use #### headings for close-reading subsections because #### is reserved for paper-card titles and is checked by paper-card-delivery validators.
- After the source-order close reading, write one integrated mechanism synthesis. Its headings may vary with the paper, but it must cover the strongest author argument, assumptions and falsifiable predictions, claim-evidence-alternative-explanation alignment, counterfactual ablation predictions, minimal necessary design, failure boundaries, and a discriminating next research question. Add user-specific transfer only at the end and only when it follows naturally from the paper.
- Validate the Paper Card placed inside the main-page
精读部分 using paper-card-delivery; run its validator when a local Markdown draft exists.
- Store figures and assets in a stable assets folder.
- Mark author claim, experimental support, inference, citation needed, and unresolved questions separately.
Paper-card content standards live in paper-card-delivery. This deep-dive skill must not duplicate or override paper-card source verification, metadata, image/caption selection, fixed bullet slots, sorting, or structural validation.
Local MinerU Extraction
For future deep dives, first create a MinerU conversion draft when a PDF is available. Use it as the source-order scaffold for the main-page 原文中译稿 and 精读部分, and for an English manuscript only when that optional artifact is explicitly requested.
- Preferred wrapper in this vault:
$WORLD_MODEL_VAULT/.tools/mineru-md.sh
- MinerU binary on this machine:
$WORLD_MODEL_VAULT_MINERU_BIN
- Verified local version:
mineru 3.3.1
- Store MinerU outputs, downloaded PDFs, supplementary/appendix PDFs, official HTML snapshots/pages, arXiv HTML, LaTeX source, and temporary figure assets under
.tools/tmp/codex/<task-slug>/; delete them after the target artifacts are written and verified successfully.
- MinerU is a conversion draft, not the authoritative final text. When an official HTML version exists, always check the MinerU draft against it before publishing target artifacts. For arXiv papers, arXiv HTML is the preferred HTML check; for non-arXiv papers, use publisher/proceedings/OpenReview/CVF HTML when available. Verify section order, paragraph continuity, equations, figures, captions, tables, appendices, citations, and references. If HTML is unavailable or incomplete, use official LaTeX source or the official PDF as the authority and record that HTML could not be used.
- If MinerU misses or corrupts formulas, figures, captions, appendices, or references, repair from official HTML/LaTeX/PDF or the official publisher source before marking the deep dive complete.
- If MinerU itself fails but the PDF is accessible, try the local wrapper again with a clean output directory, inspect the error, and then use a structured fallback such as official HTML/LaTeX, publisher HTML, Docling, Marker, PyMuPDF, or pdfplumber. MinerU failure is a workflow problem to resolve or work around, not permission to ship manuscript summaries.
Manuscript Fidelity Requirements
- The main-page
原文中译稿 must preserve paper-like citation flow. Preserve the source paper's citation style in the body: author-year forms such as (Hassan 等人,2019a) are valid and should not be forcibly converted to numeric citations. When a verified PDF URL exists, the citation text should carry the link. Numeric citations must remain bracketed when the source uses them, with the numeric PDF-link contract below applied.
- The main-page manuscript must cover the full source package: Abstract, Introduction, all numbered/named main sections, Conclusion/Discussion, appendices and/or supplementary materials when they exist, figure and table captions, algorithms, and References. If an optional English manuscript is requested, it must cover the same source package. Record any user-requested exclusions.
- Supplementary / appendix is first-class deep-dive content, not an optional add-on. Include it in the main-page Chinese manuscript; include it in an optional English manuscript when one is requested. A package that stops at the main conference PDF while a usable supplement exists is incomplete.
- Before marking complete, compare the main-page manuscript against the official source section list including appendix/supplement headings. Apply the same check to any optional English manuscript that is published.
- Figures and tables must be placed near their original reference/caption positions. Use native Feishu/Notion image blocks or stable relative Obsidian assets when reliable official image assets are available.
- For tables, verify both semantic extraction and visual fidelity. Compare complex or formula-heavy tables against an official PDF/HTML screenshot, and include that screenshot in
原文中译稿 when extraction cannot guarantee merged cells, multi-level headers, footnotes, symbols, colors, borders, or layout. MinerU is a manuscript scaffold and locator, not the sole authority for table screenshots.
- Attach captions using the selected platform's native or established representation: native image captions in Feishu/Notion, and the vault convention or meaningful alt text in Obsidian. Captions with formulas may use an immediately adjacent formula-capable block when the native caption cannot preserve TeX; preserve the exact TeX source and do not duplicate the caption.
- Figure captions are source-fidelity content. The main-page Chinese manuscript must use a complete Chinese translation of each caption. An optional English manuscript must preserve the official original caption. Do not replace captions with agent-written summaries or source-process notes.
- Table translation rule for
原文中译稿: translate table captions and table notes into Chinese; do not translate table cell content, including headers and body cells. An optional English manuscript keeps both captions and cells in the original English.
- Table screenshot support:
原文中译稿 may and should include an original table screenshot when the table's visual structure cannot be trusted after extraction. Use a crop from the official PDF or a verified official HTML rendering as the visual authority; do not treat a MinerU-generated table image or reconstructed screenshot as authoritative. Place the screenshot near the corresponding table position, preserve the official English cell content in the screenshot, and put the translated Chinese table title and table notes into the platform's native image caption. Do not duplicate the same caption or notes as adjacent body prose. If the target platform supports an editable table, retain the editable English-cell table as well; the screenshot is the visual-fidelity reference, not an excuse to drop the table entirely. If only a screenshot can preserve the table reliably, label it as an original table screenshot and record the rendering fallback in the verification notes.
- Formulas must be checked against official HTML/LaTeX/PDF and preserved in LaTeX where possible. This includes inline formulas, not only displayed equations. Do not publish pages where important equations, inline variables, losses, or symbolic expressions have collapsed into prose or lost subscripts/superscripts.
- The complete Chinese manuscript follows an official HTML/LaTeX-derived source map in original section order, one natural paragraph at a time, including figure/table positions, formula placement, body citations, References, captions, appendices/supplements, and table structure. For
pdf2zh imports, reconstruct paragraph boundaries from that source map: merge page/column fragments, split falsely fused paragraphs, remove duplicated overlap, and correct reordered columns. Chinese punctuation, paragraph length, PDF text extraction, and current platform block boundaries are anomaly detectors only, never the final authority when official HTML/LaTeX exists. Its References section remains the source-faithful bibliography with the same [n] labels and PDF-link contract derived from the source map.
- Chinese terminology must be deliberate. Translate generic technical concepts into accurate Chinese, but keep named architectures, models, methods, datasets, and official components unchanged in their source form, including names such as
Transformer and DINO. Add a short Chinese gloss on first use only when it helps comprehension; do not invent translations for official names. Do not leave dense generic English terminology untranslated in ordinary Chinese explanatory prose. After drafting the Chinese manuscript, run the dedicated terminology correction gate before marking it complete.
Reference and Citation Link Contract (mandatory)
This contract applies to the required main-page 原文中译稿 and, when explicitly requested, the optional <paper short name>|英文原文稿. The Chinese manuscript must not invent a different citation scheme.
References block (bibliography)
- In the English manuscript, do not translate any bibliography field. In the Chinese manuscript, the cited paper title may be translated for readability, including when the title was already translated by
pdf2zh-next; authors, venues, publishers, page ranges, year, DOI/arXiv strings, URLs, and all other bibliographic metadata must remain in the original language and source order.
- Format: plain text lines / paragraphs starting with bracket labels such as
[1], [2], [12]. One reference per line or paragraph.
- Title delimiters: enclose each cited paper title in Chinese quotation marks
“...” so the title is visually distinct from authors, venue, year, and other bibliographic metadata. This applies whether the title is retained in English or translated in a Chinese manuscript.
- Important references: when a reference is central to the paper's method, baseline, theoretical foundation, or the user's research context, underline only the paper title using the target platform's native underline formatting. In Notion enhanced Markdown, use
<span underline="true">...</span>; do not underline the citation number, authors, venue, year, URL, or the entire reference entry. Do not overuse this emphasis: mark only genuinely important references.
- Forbidden formats: Markdown/platform ordered lists (
1. 2. 3.), bullet lists that replace or hide the bracket numbers, renumbered citations, or Chinese-translated bibliography entries.
- PDF link suffix (mandatory when a usable PDF URL can be found): after the full reference text, append
. URL (period, space, URL, space) — or just URL if the bibliography text already ends with a period, to avoid .. URL — and then a hyperlink whose display text equals the URL string itself. Prefer an arXiv PDF URL of the form https://arxiv.org/pdf/<id> (with or without version, matching the cited work). If no arXiv PDF exists, use the best open PDF (OpenReview, CVF, PMLR, publisher OA, project page) in the same . URL [url](url) form. If no PDF can be verified after search, keep the entry without a fake link and mark PDF link: not found in the verification log—do not invent URLs.
- Canonical References line example (Markdown):
[1] Author A, Author B. “Paper title.” Conference/Journal, year. URL [https://arxiv.org/pdf/2401.12345](https://arxiv.org/pdf/2401.12345)
- Important-reference example (Notion):
[12] Author A, Author B. “<span underline="true">Paper title</span>.” Conference/Journal, year. URL [https://arxiv.org/pdf/2401.12345](https://arxiv.org/pdf/2401.12345)
- On Feishu/Notion, render the same structure:
[n] plain label + original English bibliography text + . URL (or URL after an existing period) + clickable URL whose visible text is the full URL.
Body citations
- Preserve every in-text citation that corresponds to the bibliography in the source paper's native style. Author-year citations such as
(Hassan 等人,2019a) are valid in the Chinese manuscript when that is the source style; an optional English manuscript keeps (Hassan et al., 2019a). When a verified PDF URL exists, link the citation text itself. Numeric citations must remain bracketed and in source positions.
- Link only the number, never the citation brackets. The canonical body-citation form is
[[1](https://arxiv.org/pdf/2401.12345)]: the inner 1 is the hyperlink text, while the outer [ and ] are ordinary plain-text citation brackets. The link range must not include either bracket. Do not use [[1]](https://arxiv.org/pdf/2401.12345) or any platform form that makes [1] the hyperlink text. Target the same PDF URL recorded for that [n] in References:[[1](https://arxiv.org/pdf/2401.12345)]
Multi-cite example:[[1](https://arxiv.org/pdf/2401.12345), [3](https://arxiv.org/pdf/2305.67890)]
- Platform read-back check: after writing to Feishu or Notion, inspect the rendered link range. It must show a plain outer bracket before and after a linked numeral, visually
[1]; selecting the link must select only 1, not [1]. On Notion, never write bare [[n](url)]: escape the outer bracket (\[[n](url)] / multi \[[n](url), [m](url)]) via notion-doc-workflow/scripts/prepare-notion-citation-markdown.py before Markdown write—single or multi, the first cite is always the corrupted one. After write, still run notion-doc-workflow/scripts/fix-notion-citation-rich-text.py <page> and require --check-only to exit 0 before delivery.
- The body link target for numeric
[n] citations must be identical to
…(truncated)
1---2name: paper-deep-dive3description: Canonical single-paper deep-dive delivery standard, with Notion as the durable target. Use whenever the user mentions deep dive, 深读, 详细解析, detailed-read, dive into a paper, full paper reading, 原文中译稿, or asks to audit/repair a deep-dive package. Produces one self-contained Notion main page whose primary body is the complete 原文中译稿 and whose 精读部分 contains the Paper Card, editable 论文解析树, source-order close reading, and mechanism synthesis. English-manuscript and standalone close-reading child pages are optional only when explicitly requested.4---56# Paper Deep Dive78Short description: PDF to deep paper notes workflow. Historical source: Feishu wiki AI Research Skills; current durable delivery is Notion-first.910Use this skill when the user wants a complete paper reading workflow from PDF to stable notes, paper card, assets, and shareable summaries.1112## Single Source of Truth1314This skill is the only canonical delivery standard for single-paper deep dives. Other skills may trigger, route, write to Feishu/Notion/Obsidian, preserve Chinese wording, or distill the finished deep dive into a wiki, but they must not define a second deep-dive structure or mark a package complete under weaker rules.1516When any user request says `deep dive`, `深读`, `详细解析`, `detailed-read`, `dive into`, `精读这篇论文`, or requests a complete `原文中译稿`, treat it as this full workflow unless the user explicitly asks for a lighter artifact such as quick summary, paper card only, partial translation, or English manuscript only. A lighter artifact must be labeled as such and must not be called a compliant deep dive.1718The completion standard is product-level, not effort-level. A good summary or a partial translation is still incomplete. A package is compliant when the main page contains the complete Chinese manuscript and its required 精读部分, and figures/captions, formulas, references, hierarchy, and target-platform read-back verification all pass the delivery gate below. Child pages are not part of the default completion gate.1920## Canonical Platform Delivery Gate2122Also use [`research-doc-workflow`](../research-doc-workflow/SKILL.md) and [`notion-doc-workflow`](../notion-doc-workflow/SKILL.md) for new deep dives. Notion is the canonical durable target because Feishu storage is limited. Keep one Markdown-first manuscript and close-reading structure; only Notion hierarchy, native equations/captions, uploads, editable trees, citation chrome, and write verification need platform adaptation. Do not create or maintain a parallel Feishu deep-dive copy unless the user explicitly requests a legacy repair or migration.2324## Canonical Paper Card Gate2526When this workflow creates or modifies a paper card, also use [`paper-card-delivery`](../paper-card-delivery/SKILL.md). That skill is the canonical standard for official-source verification, fixed card format, figure/caption selection, and structural validation. Do not finalize the card from this deep-dive skill alone.2728## Canonical Chinese Technical Writing Gate2930For `原文中译稿`, `完整中文译稿`, `精读稿`, Chinese figure captions, Chinese table captions/notes, main-entry summaries, and Chinese paper-card prose, also use [`chinese-technical-writing`](../chinese-technical-writing/SKILL.md). Preserve official English source text in `英文原文稿`, formulas, named architectures/models/methods such as `Transformer` and `DINO`, dataset names, symbols, and table cell content. Translate generic technical concepts into Chinese, but do not translate official names or force them into awkward Chinese names.3132## Borrowed Method Layer3334This is the user's canonical deep-dive workflow. It may borrow useful reading methods from downloaded skills, but those skills do not own the final package semantics or target-platform representation.3536- From `nature-reader`, borrow the source-map-first method: build stable source37 block IDs such as `S001`, `F001`, `T001`, preserve original / Chinese38 correspondence, keep figure/table captions attached to the relevant text, and39 record uncertainty instead of guessing.40- Do not publish the external `nature-reader` artifact contract as-is. A lone `paper.md`, `source_map.json`, `translation_notes.md`, or bilingual reader does not replace the required self-contained main page. These files may be intermediate material or part of an explicitly requested Obsidian package.41- Use the source map as an internal scaffold for the main-page `原文中译稿` and source-grounded `精读部分`; use it for an `英文原文稿` only when the user explicitly requests that optional artifact. The final deliverable must use the selected platform's native images/captions, formulas, hierarchy, and read-back verification through `research-doc-workflow`.42- In the main-page `精读部分`, short bilingual source snippets or block IDs may be included when they clarify a key claim, equation, or figure, but this section remains analytical material rather than a second translation.4344## Mechanism Interrogation Gate4546Treat deep reading as reconstruction followed by audit. First recover the strongest version of the authors' logic; only then test whether the problem is important, the assumptions are defensible, the design addresses the stated bottleneck, the evidence excludes relevant alternatives, and the conclusion stays within the evidence boundary. Do not confuse skepticism with automatic rejection.4748Before calling the analytical close reading complete, establish all of the following from the source package:4950- the concrete prior-method bottleneck and its proposed causal explanation;51- the key assumptions and the falsifiable predictions they imply;52- a design-to-mechanism chain: `design -> changed information/constraint/optimization -> predicted effect`;53- the decisive experiment, control, or ablation for each central mechanism claim;54- plausible alternative explanations that the experiments do or do not rule out;55- counterfactual predictions for removing, replacing, or simplifying a module;56- the data, scene, supervision, or optimization conditions under which the method should fail;57- the smallest credible alternative design and the next question that would discriminate between explanations.5859Keep `作者声称`, `实验支持`, `我们的推断`, and `尚未验证` distinguishable in analytical notes. These are evidence statuses, not mandatory repeated headings. Do not inject this analysis into the source-faithful `原文中译稿`; it belongs in the main-page `精读部分`, including the editable tree and analytical close reading.6061## When To Use6263- Reading a new paper deeply rather than only summarizing it.64- Converting the verified paper source into one self-contained main page: complete faithful Chinese translation plus a clearly separated 精读部分 containing the paper card, editable paper-analysis tree, source-order close reading, and mechanism synthesis.65- Preparing Notion paper pages/notes, or repairing an existing legacy page in another platform when explicitly requested.66- Auditing whether figures, claims, assets, and citations are complete.6768If the user says `deep dive`, `深读`, `详细解析`, `dive into`, or asks to deeply read a single paper, use this full workflow by default. Do not downgrade it to a quick summary, paper card only, or close-reading note only unless the user explicitly asks for a lighter output.6970## Non-Negotiable Deliverable7172For papers with an accessible official PDF or full-paper HTML, the sole mandatory durable artifact is the Notion main page. Its primary body is the complete faithful `原文中译稿` in source order, followed by a clearly separated, embedded section titled `精读稿`. That section contains the Paper Card, editable `论文解析树`, source-order analytical close reading, and integrated mechanism synthesis.7374`<paper short name>|英文原文稿` and a standalone `<paper short name>|精读稿` child page are optional artifacts. Do not create either by default. Create them only when the user explicitly requests them or supplies an existing hierarchy that must be preserved during migration/repair. Their absence is never a deep-dive completion failure.7576If the PDF can be downloaded or viewed, assume the complete Chinese manuscript can be produced by MinerU extraction plus official HTML / LaTeX / PDF verification. Do not use context length, page length, one-turn time, target-page size, translation workload, or "current tool path" as reasons to downgrade it into a section summary, structured outline, selected excerpts, or partial translation. Chunk the paper by sections, append incrementally, and continue until the main page is complete.7778Only three conditions justify not producing the complete main-page Chinese manuscript: the full paper source is inaccessible, reproduction is blocked by a clear licensing/copyright constraint, or the user explicitly asks for a lighter / partial artifact. Otherwise, an incomplete `原文中译稿` is work in progress, not a compliant deep dive.7980## Source Acquisition Priority8182Before running MinerU or writing any target page/note, build the complete source package, including every official or source-derived HTML version that can be found. Search for arXiv HTML (`https://arxiv.org/html/<id>` and versioned variants), publisher/proceedings HTML, OpenReview/forum HTML, CVF/open-access HTML, and project-page paper HTML when present. HTML is not optional source decoration: it is often the best source for section hierarchy, MathML / TeX annotations, figure/table nodes, captions, references, and supplementary links.8384For papers that have an arXiv version, search and inspect the arXiv landing page, arXiv PDF, arXiv HTML, and LaTeX source before using a conference or publisher PDF as the main extraction source. Prefer arXiv HTML / LaTeX for structure, formulas, captions, references, and appendix discovery; prefer arXiv PDF for page layout, figure placement, and visual cross-checking. The reason is practical: many official conference PDFs omit appendices or supplementary sections, while arXiv often preserves the fuller manuscript.8586Use venue/publisher pages and their official article/proceedings HTML for official metadata, acceptance venue, project/code links, supplement links, and cross-checking, but do not assume the venue PDF is the complete manuscript. If the only PDF initially found is from CVF, ACM, IEEE, Springer, PMLR, OpenReview, NeurIPS, a conference proceedings site, or a publisher landing page, explicitly search for separate `supplementary`, `appendix`, `supp`, `supplemental material`, `additional material`, `SM`, or `PDF supplementary` files on the same page, the official HTML page, the proceedings page, OpenReview, the project page, author/lab page, and arXiv.8788When HTML and PDF disagree, treat official HTML / LaTeX as the first authority for source text, section order, formulas, captions, and references when it clearly preserves the paper source; cross-check figure placement, page layout, missing appendices, and visual assets against the PDF and supplement. Record meaningful discrepancies instead of silently choosing one source.8990If a separate appendix/supplementary PDF exists, it is part of the deep-dive source package unless the user explicitly excludes it. Download it alongside the main PDF, include its sections, figures, tables, formulas, captions, algorithms, and references in the source map, and reflect it in the main-page `原文中译稿`. If an optional English manuscript is explicitly requested, include the same source coverage there. A deep dive based only on a conference main PDF is incomplete when a separate supplement exists and has not been inspected.9192**Supplementary / appendix PDF extraction priority (mandatory):** when a separate supplementary or appendix PDF has no usable official HTML (typical for project-page or ACM/IEEE “Supplementary Material” PDFs), the required first conversion path is MinerU via `.tools/mineru-md.sh` (same Local MinerU Extraction rules as the main PDF). Do not use PyMuPDF / pdfplumber / raw `get_text` / page screenshots as the primary manuscript scaffold for that supplement. After MinerU, repair formulas/captions/tables against the PDF (and against HTML/LaTeX only if a supplement HTML/LaTeX source actually exists). PyMuPDF and similar tools are allowed only as a fallback after MinerU fails, or for narrow visual checks (page count, figure crop verification), not as the default text path for supplements.9394Record the source package in task notes and compact source metadata: arXiv ID/version when available, arXiv PDF status, arXiv HTML URL/status, LaTeX source status, publisher/proceedings/OpenReview/CVF HTML URL/status, venue/publisher PDF status, supplementary/appendix PDF status, project/code links, and extraction date. If no HTML version is found, state the searched routes briefly; do not write `HTML not checked`. If no supplement is found, state the searched routes briefly; do not write `supplement not checked`.9596## Workflow9798### Imported PDF Translation Repair Path99100When the user provides an arXiv PDF or paper title and asks to download, run101`pdf2zh`, and import the Chinese PDF into Notion, also use102[`paper-pdf2zh-notion-import`](../paper-pdf2zh-notion-import/SKILL.md). That103skill owns the direct Notion desktop import, title-based PDF renaming, local104artifact cleanup, and the mandatory two-round105repair/read-back cycle. The imported PDF is always draft material; a successful106Notion upload is not a completed `原文中译稿`.107108When the user has already translated a paper PDF with `pdf2zh-next` and manually imported the resulting PDF into Notion or Feishu, treat the imported page as a **draft translation**, not as a completed `原文中译稿`. This path is an accepted entry point and does not require re-running PDF translation, but it must pass the same source-fidelity and platform verification gates before delivery.109110**One-line workflow trigger:** `帮我 pdf2zh 并导入 Notion` is sufficient to invoke the full import workflow. Do not ask the user to restate the repair rules. Resolve the latest arXiv PDF, run the local conversion, import the Chinese PDF directly into the supplied Notion parent, then perform the two-round repair/read-back cycle below.111112**Imported-page delivery rule:** The existing imported Chinese page is the primary deliverable, matching the standard deep-dive structure. Do **not** require creation of `英文原文稿` or standalone `精读稿` child pages unless the user explicitly asks for them. Verify the Chinese manuscript against the latest official PDF and available HTML/LaTeX, repair source-order structure, formulas, figures and native captions, tables, appendices, references, and body citations, add or repair the required main-page `精读部分`, then perform target-platform read-back verification.113114- **A. Identify the source.** Locate the imported page and source PDF. Preserve imported figures, tables, and page order as draft material, then compare them against the source PDF and, when available, official HTML / LaTeX.115- **B. Fix identity and navigation.** Rename the page to the verified Chinese paper title only. Keep the official English title in the opening block. Add the latest arXiv PDF, Project Page, and Code links above the abstract when available. Do not add `英文原文稿` or standalone `精读稿` child links unless explicitly requested; never infer URLs from a filename.116- **C. Repair structure.** Restore heading levels from the paper's numbered hierarchy, reconnect PDF-split paragraphs, remove duplicate headers/footers and conversion artifacts, and keep figures/tables near their source positions.117- **C1. Normalize numbered headings at every depth.** Use the numeric prefix to determine hierarchy, not the importer’s visual level: a one-part prefix such as `N.` (`3.`, `4.`) is a top-level section; a two-part prefix such as `N.M.` (`3.1.`, `4.2.`) is its subsection; a three-part prefix such as `N.M.K.` (`3.1.1.`, `4.2.1.`) is its sub-subsection; continue the same rule for deeper prefixes. Normalize full-width heading punctuation `.` (U+FF0E) to the ASCII period `.` before parsing, so `4.2.方法` becomes `4.2. 方法`. In ordinary structural text, normalize full-width slash `/` (U+FF0F) to `/`, full-width hyphen-minus `-` (U+FF0D) to `-`, and full-width brackets `[]` (U+FF3B/U+FF3D) to `[]` when they are citation, list, or link delimiters. Protect LaTeX, code, URLs, file paths, and existing backslash-escaped sequences before this cleanup, then restore them exactly; never perform a blind global replacement. Do not replace unrelated Chinese punctuation in ordinary prose. Default to `## 3. Section title`, `### 3.1. Subsection title`, and `#### 3.1.1. Sub-subsection title`. If the source consistently omits punctuation (`3 Title`, `3.1 Subtitle`), preserve that style for the whole manuscript; never mix punctuated and unpunctuated forms or place a deeper numeric prefix above its parent.118- **D. Repair formulas.** Restore inline formulas as native inline equations or exact `$...$` LaTeX, restore display equations and numbering, and sample early, middle, formula-heavy, and appendix sections.119- **E. Repair figures and tables.** Convert figure captions to native caption fields and keep the complete translated source caption without a duplicate paragraph or agent-written summary. Preserve English table cells, translate only table titles/notes, and use an original PDF/HTML screenshot when table layout is unreliable.120- **E0. Caption parser check.** Before writing a native Notion/Feishu image caption, verify that escaped citation delimiters such as `\[33\]` will not be parsed as block syntax. If the caption parser cannot preserve them, convert only those citation markers inside the caption to ordinary visible parentheses while keeping body citations and reference numbering unchanged; then verify that the caption is stored on the image block and no duplicate caption paragraph remains.121- **E1. Repair imported table corruption.** Inspect every imported table block and its surrounding text for OCR/PDF conversion residue: duplicated Markdown pipe tables after a native table, broken rows or columns, repeated cell fragments, garbled characters, malformed separators, and captions fused to table data. Keep one authoritative editable table, reconstruct rows/cells from the official PDF/HTML when the extracted structure is reliable, and otherwise add an official PDF/HTML table screenshot as the visual authority. Remove duplicate pseudo-tables and keep exactly one translated table title/note attached to the table.122- **E2. Author contacts and resource links.** Author email addresses must be written as ordinary visible text, not intentionally wrapped in Markdown links, `mailto:` links, or code formatting. Notion may auto-link a bare email during rendering; do not add an explicit link or change its visible text to code merely to fight that platform behavior. **Only in the opening resource block of a deep-dive Chinese manuscript**, omit the heading `来源` and display each link's URL as its link text, for example `[https://arxiv.org/pdf/<id>](https://arxiv.org/pdf/<id>)`. This URL-as-label rule does not apply to English manuscripts, paper cards, or general research documents.123- **F. Repair citations.** A Chinese-manuscript reference **title may be translated**, but authors, venue, publisher, year, volume/issue, pages, DOI/arXiv identifiers, URLs, and other bibliographic metadata remain in the source language. Keep `[n]` labels, one reference per paragraph, and restore verified PDF URLs. Body citation links must use the same URL map.124- **G. Verify completion.** Run the full main-page completion gate and record unresolved figure, table, formula, reference, URL, hierarchy, Paper Card, editable-tree, or close-reading issues. A manually imported page is complete only after repair and fetch/read-back verification; the absence of optional child artifacts is not a failure.125126**Two-round repair is mandatory for imported pages.** Round 1 fixes structure and source fidelity against HTML/LaTeX/PDF. Round 2 starts from a fresh Notion read-back and independently audits formulas, inline math, figure/caption placement, table integrity, references, body citation links, appendices, and conversion residue. A first-pass upload or a single visual scan is never a completed delivery.1271281. Capture source metadata and source package inventory: title, authors, year, venue, DOI/arXiv, arXiv PDF URL/status, arXiv HTML URL/status, LaTeX source status, publisher/proceedings/OpenReview/CVF HTML URL/status, venue/publisher PDF URL, supplementary/appendix PDF URL(s), project/code links, local source paths, and extraction date.1292. Extract or parse the complete source package to inspectable Markdown when tooling is available; preserve figure references and equation context. On this machine, use MinerU as the default PDF-to-Markdown path before building deep-dive artifacts, but parse official HTML / LaTeX first when available for source structure, formulas, captions, references, and appendix coverage. Parse the arXiv/full manuscript first when available, then parse any separate supplement/appendix PDF with the same priority: official supplement HTML/LaTeX if present, otherwise MinerU on the supplement PDF before any other PDF text extractor.1303. Check the MinerU conversion draft against official HTML whenever HTML exists, especially arXiv HTML for arXiv papers and publisher/proceedings HTML for non-arXiv papers. This check is mandatory, not optional. Repair section order, paragraph continuity, formulas, figures, tables, captions, appendices, body citations, and references before publishing.1314. Build a source map inspired by `nature-reader`: stable block IDs for body text, figures, tables, captions, equations, appendices, and references; page / section location; extraction confidence; and links between first figure/table mention and the visual asset.1325. Build a verified English source scaffold internally from official HTML/LaTeX/PDF. It must be complete enough to support paragraph-level translation and audit, but it is a working artifact, not a required reader-facing page. If the user explicitly requests `<paper short name>|英文原文稿`, publish the complete original English text in source order and verify it independently.1336. **English source correction gate (mandatory before Chinese translation):** for arXiv papers, re-open arXiv HTML and correct the internal English scaffold against it. Check section order, paragraph continuity, formulas, figure/table captions, appendix/supplement coverage, body citations, and References. Prefer arXiv HTML for text/structure/formulas and use PDF for layout/visual cross-check. For non-arXiv papers, use the best official HTML; if none exists, correct against LaTeX/PDF and record that HTML correction was impossible. This gate applies even when no English child page is published.1347. Create the complete faithful Chinese manuscript directly in the main page from the corrected source scaffold. It must preserve section hierarchy, paragraph correspondence, formulas, figure/table positions, citations, captions, references, appendices/supplements, and layout structure as much as the target editor allows. Translate the paper body, figure captions, table captions/notes (表注), appendix/supplement prose, and explanatory text into Chinese, but keep table cell content in the original English and keep References / bibliography entries source-faithful with the `[n] … . URL [url](url)` PDF-link contract. A partial translation is allowed only as a clearly marked WIP state.1358. **Chinese terminology correction gate (mandatory after the Chinese manuscript draft exists):** do a dedicated second pass over `原文中译稿` for terminology only. Verify key method/model/dataset/loss/module terms are consistent; keep named architectures, models, methods, datasets, and official components such as `Transformer` and `DINO` in their official English form; add a Chinese gloss only when it improves comprehension; and translate generic technical concepts instead of leaving avoidable English phrase islands. Fix inconsistent renderings of the same term across sections. Do not mark the Chinese manuscript complete until this terminology pass is done.1368b. **Reference / citation verification gate (mandatory for the main manuscript):** build a single `[n] → PDF URL` map (prefer arXiv PDF), apply it to References (`. URL [url](url)`) and body (`[[n](url)]`), then verify URL correctness and cross-consistency as specified in Reference and Citation Link Contract. If an optional English manuscript is published, apply and verify the same map there. For an imported PDF2ZH Chinese manuscript, the cited paper title may be translated, but all other bibliographic fields must remain source-faithful.1379. Create or update the main reader-facing deep-dive page. The primary body is the complete `原文中译稿` and the default reading surface. Its opening must follow a paper-like title block before the abstract: official English title, Chinese title, original English author list and affiliations, then separate verified links for the latest arXiv PDF, Project Page, and Code. Continue with the source-faithful abstract and manuscript in normal paper order. After the full manuscript, add a clearly separated `精读部分`; do not interleave analysis with the translation. Do not create or link child artifacts unless explicitly requested.13810. Create an editable `论文解析树` that follows the paper's actual reasoning: problem -> concrete bottleneck -> key assumption -> design/mechanism -> changed information or constraint -> predicted effect -> decisive evidence -> boundary. Make the information-flow view (what passes between modules) and the causal-chain view (why the design should change the result) distinguishable. Add losses/training, datasets/evaluation, limitations, and user research implications where they clarify this logic rather than as disconnected inventory branches. Use a native Feishu mind map for Feishu, a structured page/database or supported embedded artifact for Notion, and Mermaid/Canvas plus a searchable linked outline for Obsidian. Do not substitute a static screenshot when an editable representation is available.13911. In the main-page `精读部分`, create the Paper Card first, then the editable `论文解析树`, then a source-order analytical close reading and integrated mechanism synthesis. Follow the paper's own section order and local context: Abstract / Introduction, numbered sections, named subsections, conclusion, then appendices or supplementary material. For each part, explain which claim it advances, why that step is needed, what mechanism or evidence is introduced, and what remains unresolved. Do not insert a repeated per-section heading or paragraph such as "what this means for my world-model research" / "对你的 world model 研究意味着什么". Put user-specific implications and future project ideas only in the final synthesis. This is interpretation, not part of the source-faithful translation.140 - Inside `精读部分`, use `###` or lower-impact paragraph/list structure for source-order close-reading subsections. Do not use `####` headings for close-reading subsections because `####` is reserved for paper-card titles and is checked by `paper-card-delivery` validators.14112. After the source-order close reading, write one integrated mechanism synthesis. Its headings may vary with the paper, but it must cover the strongest author argument, assumptions and falsifiable predictions, claim-evidence-alternative-explanation alignment, counterfactual ablation predictions, minimal necessary design, failure boundaries, and a discriminating next research question. Add user-specific transfer only at the end and only when it follows naturally from the paper.14213. Validate the Paper Card placed inside the main-page `精读部分` using [`paper-card-delivery`](../paper-card-delivery/SKILL.md); run its validator when a local Markdown draft exists.14314. Store figures and assets in a stable assets folder.14415. Mark author claim, experimental support, inference, citation needed, and unresolved questions separately.145146Paper-card content standards live in [`paper-card-delivery`](../paper-card-delivery/SKILL.md). This deep-dive skill must not duplicate or override paper-card source verification, metadata, image/caption selection, fixed bullet slots, sorting, or structural validation.147148## Local MinerU Extraction149150For future deep dives, first create a MinerU conversion draft when a PDF is available. Use it as the source-order scaffold for the main-page `原文中译稿` and `精读部分`, and for an English manuscript only when that optional artifact is explicitly requested.151152- Preferred wrapper in this vault: `$WORLD_MODEL_VAULT/.tools/mineru-md.sh`153- MinerU binary on this machine: `$WORLD_MODEL_VAULT_MINERU_BIN`154- Verified local version: `mineru 3.3.1`155- Store MinerU outputs, downloaded PDFs, supplementary/appendix PDFs, official HTML snapshots/pages, arXiv HTML, LaTeX source, and temporary figure assets under `.tools/tmp/codex/<task-slug>/`; delete them after the target artifacts are written and verified successfully.156- MinerU is a conversion draft, not the authoritative final text. When an official HTML version exists, always check the MinerU draft against it before publishing target artifacts. For arXiv papers, arXiv HTML is the preferred HTML check; for non-arXiv papers, use publisher/proceedings/OpenReview/CVF HTML when available. Verify section order, paragraph continuity, equations, figures, captions, tables, appendices, citations, and references. If HTML is unavailable or incomplete, use official LaTeX source or the official PDF as the authority and record that HTML could not be used.157- If MinerU misses or corrupts formulas, figures, captions, appendices, or references, repair from official HTML/LaTeX/PDF or the official publisher source before marking the deep dive complete.158- If MinerU itself fails but the PDF is accessible, try the local wrapper again with a clean output directory, inspect the error, and then use a structured fallback such as official HTML/LaTeX, publisher HTML, Docling, Marker, PyMuPDF, or pdfplumber. MinerU failure is a workflow problem to resolve or work around, not permission to ship manuscript summaries.159160## Manuscript Fidelity Requirements161162- The main-page `原文中译稿` must preserve paper-like citation flow. Preserve the source paper's citation style in the body: author-year forms such as `(Hassan 等人,2019a)` are valid and should not be forcibly converted to numeric citations. When a verified PDF URL exists, the citation text should carry the link. Numeric citations must remain bracketed when the source uses them, with the numeric PDF-link contract below applied.163- The main-page manuscript must cover the full source package: Abstract, Introduction, all numbered/named main sections, Conclusion/Discussion, **appendices and/or supplementary materials when they exist**, figure and table captions, algorithms, and References. If an optional English manuscript is requested, it must cover the same source package. Record any user-requested exclusions.164- **Supplementary / appendix is first-class deep-dive content**, not an optional add-on. Include it in the main-page Chinese manuscript; include it in an optional English manuscript when one is requested. A package that stops at the main conference PDF while a usable supplement exists is incomplete.165- Before marking complete, compare the main-page manuscript against the official source section list including appendix/supplement headings. Apply the same check to any optional English manuscript that is published.166- Figures and tables must be placed near their original reference/caption positions. Use native Feishu/Notion image blocks or stable relative Obsidian assets when reliable official image assets are available.167- For tables, verify both semantic extraction and visual fidelity. Compare complex or formula-heavy tables against an official PDF/HTML screenshot, and include that screenshot in `原文中译稿` when extraction cannot guarantee merged cells, multi-level headers, footnotes, symbols, colors, borders, or layout. MinerU is a manuscript scaffold and locator, not the sole authority for table screenshots.168- Attach captions using the selected platform's native or established representation: native image captions in Feishu/Notion, and the vault convention or meaningful alt text in Obsidian. Captions with formulas may use an immediately adjacent formula-capable block when the native caption cannot preserve TeX; preserve the exact TeX source and do not duplicate the caption.169- Figure captions are source-fidelity content. The main-page Chinese manuscript must use a complete Chinese translation of each caption. An optional English manuscript must preserve the official original caption. Do not replace captions with agent-written summaries or source-process notes.170- Table translation rule for `原文中译稿`: translate table captions and table notes into Chinese; do **not** translate table cell content, including headers and body cells. An optional English manuscript keeps both captions and cells in the original English.171- Table screenshot support: `原文中译稿` may and should include an original table screenshot when the table's visual structure cannot be trusted after extraction. Use a crop from the official PDF or a verified official HTML rendering as the visual authority; do not treat a MinerU-generated table image or reconstructed screenshot as authoritative. Place the screenshot near the corresponding table position, preserve the official English cell content in the screenshot, and put the translated Chinese table title and table notes into the platform's native image caption. Do not duplicate the same caption or notes as adjacent body prose. If the target platform supports an editable table, retain the editable English-cell table as well; the screenshot is the visual-fidelity reference, not an excuse to drop the table entirely. If only a screenshot can preserve the table reliably, label it as an original table screenshot and record the rendering fallback in the verification notes.172- Formulas must be checked against official HTML/LaTeX/PDF and preserved in LaTeX where possible. This includes inline formulas, not only displayed equations. Do not publish pages where important equations, inline variables, losses, or symbolic expressions have collapsed into prose or lost subscripts/superscripts.173- The complete Chinese manuscript follows an official HTML/LaTeX-derived source map in original section order, one natural paragraph at a time, including figure/table positions, formula placement, body citations, References, captions, appendices/supplements, and table structure. For `pdf2zh` imports, reconstruct paragraph boundaries from that source map: merge page/column fragments, split falsely fused paragraphs, remove duplicated overlap, and correct reordered columns. Chinese punctuation, paragraph length, PDF text extraction, and current platform block boundaries are anomaly detectors only, never the final authority when official HTML/LaTeX exists. Its References section remains the source-faithful bibliography with the same `[n]` labels and PDF-link contract derived from the source map.174- Chinese terminology must be deliberate. Translate generic technical concepts into accurate Chinese, but keep named architectures, models, methods, datasets, and official components unchanged in their source form, including names such as `Transformer` and `DINO`. Add a short Chinese gloss on first use only when it helps comprehension; do not invent translations for official names. Do not leave dense generic English terminology untranslated in ordinary Chinese explanatory prose. After drafting the Chinese manuscript, run the dedicated terminology correction gate before marking it complete.175176## Reference and Citation Link Contract (mandatory)177178This contract applies to the required main-page `原文中译稿` and, when explicitly requested, the optional `<paper short name>|英文原文稿`. The Chinese manuscript must not invent a different citation scheme.179180### References block (bibliography)181182- In the English manuscript, do not translate any bibliography field. In the Chinese manuscript, the cited paper title may be translated for readability, including when the title was already translated by `pdf2zh-next`; authors, venues, publishers, page ranges, year, DOI/arXiv strings, URLs, and all other bibliographic metadata must remain in the original language and source order.183- **Format**: plain text lines / paragraphs starting with bracket labels such as `[1]`, `[2]`, `[12]`. One reference per line or paragraph.184- **Title delimiters**: enclose each cited paper title in Chinese quotation marks `“...”` so the title is visually distinct from authors, venue, year, and other bibliographic metadata. This applies whether the title is retained in English or translated in a Chinese manuscript.185- **Important references**: when a reference is central to the paper's method, baseline, theoretical foundation, or the user's research context, underline **only the paper title** using the target platform's native underline formatting. In Notion enhanced Markdown, use `<span underline="true">...</span>`; do not underline the citation number, authors, venue, year, URL, or the entire reference entry. Do not overuse this emphasis: mark only genuinely important references.186- **Forbidden formats**: Markdown/platform ordered lists (`1.` `2.` `3.`), bullet lists that replace or hide the bracket numbers, renumbered citations, or Chinese-translated bibliography entries.187- **PDF link suffix (mandatory when a usable PDF URL can be found):** after the full reference text, append `. URL ` (period, space, `URL`, space) — or just ` URL ` if the bibliography text already ends with a period, to avoid `.. URL` — and then a hyperlink whose **display text equals the URL string itself**. Prefer an **arXiv PDF** URL of the form `https://arxiv.org/pdf/<id>` (with or without version, matching the cited work). If no arXiv PDF exists, use the best open PDF (OpenReview, CVF, PMLR, publisher OA, project page) in the same `. URL [url](url)` form. If no PDF can be verified after search, keep the entry without a fake link and mark `PDF link: not found` in the verification log—do not invent URLs.188- **Canonical References line example (Markdown):**189 ```markdown190 [1] Author A, Author B. “Paper title.” Conference/Journal, year. URL [https://arxiv.org/pdf/2401.12345](https://arxiv.org/pdf/2401.12345)191 ```192- **Important-reference example (Notion):**193 ```markdown194 [12] Author A, Author B. “<span underline="true">Paper title</span>.” Conference/Journal, year. URL [https://arxiv.org/pdf/2401.12345](https://arxiv.org/pdf/2401.12345)195 ```196- On Feishu/Notion, render the same structure: `[n]` plain label + original English bibliography text + `. URL ` (or ` URL ` after an existing period) + clickable URL whose visible text is the full URL.197198### Body citations199200- Preserve every in-text citation that corresponds to the bibliography in the source paper's native style. Author-year citations such as `(Hassan 等人,2019a)` are valid in the Chinese manuscript when that is the source style; an optional English manuscript keeps `(Hassan et al., 2019a)`. When a verified PDF URL exists, link the citation text itself. Numeric citations must remain bracketed and in source positions.201- **Link only the number, never the citation brackets.** The canonical body-citation form is `[[1](https://arxiv.org/pdf/2401.12345)]`: the inner `1` is the hyperlink text, while the outer `[` and `]` are ordinary plain-text citation brackets. The link range must not include either bracket. Do not use `[[1]](https://arxiv.org/pdf/2401.12345)` or any platform form that makes `[1]` the hyperlink text. Target the **same PDF URL** recorded for that `[n]` in References:202 ```markdown203 [[1](https://arxiv.org/pdf/2401.12345)]204 ```205 Multi-cite example:206 ```markdown207 [[1](https://arxiv.org/pdf/2401.12345), [3](https://arxiv.org/pdf/2305.67890)]208 ```209- **Platform read-back check:** after writing to Feishu or Notion, inspect the rendered link range. It must show a plain outer bracket before and after a linked numeral, visually `[1]`; selecting the link must select only `1`, not `[1]`. **On Notion**, never write bare `[[n](url)]`: escape the outer bracket (`\[[n](url)]` / multi `\[[n](url), [m](url)]`) via `notion-doc-workflow/scripts/prepare-notion-citation-markdown.py` before Markdown write—single or multi, the first cite is always the corrupted one. After write, still run `notion-doc-workflow/scripts/fix-notion-citation-rich-text.py <page>` and require `--check-only` to exit 0 before delivery.210- The body link target for numeric `[n]` citations must be **identical** to 211212…(truncated)