ArXiv Paper Processor
Use this skill for per-paper manual summarization, with optional batch artifact download.
- Single-paper mode: process one paper directory (e.g.
<run_dir>/<arxiv_id>/).
- Batch predownload mode: process many paper directories under one run dir before writing summaries.
Language Parameter
- Use a workflow language parameter (for example
English or Chinese) and apply it manually.
- The per-paper
summary.md must be written in the selected language.
- If download scripts are called directly, pass
--language <LANG> for traceability.
Core Principle
Scripts only fetch artifacts. The model performs reading and writing.
Non-negotiable Constraint
- Do not generate
summary.md by script-based snippet extraction, regex harvesting, or template autofill.
- Do not use Python/shell scripts to auto-compose section text from abstract/introduction fragments.
- Scripts in this skill are only for artifact download (
source/pdf) and trace logs.
- The final
summary.md must come from model-side reading and synthesis of the paper content.
Optional Batch Artifact Download (Many Papers)
Use this first when Stage B has many papers:
python3 scripts/download_papers_batch.py \
--run-dir /path/to/run \
--artifact source_then_pdf \
--max-workers 3 \
--min-interval-sec 5 \
--language English
Key behavior:
- Supports
--artifact source, --artifact pdf, or --artifact source_then_pdf (default).
- Supports concurrency (
--max-workers) and safe throttling/retry (--min-interval-sec, retry args).
- Uses run-local throttle state by default (
<run_dir>/.runtime/arxiv_download_state.json) to reduce 429 risk.
- Skips papers that already have usable
source/source_extract/*.tex or existing source/paper.pdf (unless --force).
- Resume-friendly: if a paper already has a completed
summary.md, you can skip that paper's summary-writing step.
- Writes batch log to
<run_dir>/download_batch_log.json by default.
Step 1: Download Source (Preferred)
python3 scripts/download_arxiv_source.py \
--paper-dir /path/to/run/2602.00528 \
--language English
This writes:
source/source_bundle.bin
source/source_extract/
source/download_source_log.json
If usable source already exists and --force is not set, the script reuses local artifacts.
Step 2: If Needed, Download PDF
python3 scripts/download_arxiv_pdf.py \
--paper-dir /path/to/run/2602.00528 \
--language English
This writes:
source/paper.pdf
source/download_pdf_log.json
If PDF already exists and --force is not set, the script reuses local artifacts.
Step 3: Model Reads and Summarizes
- If
summary.md already exists and follows the required format, skip this paper and mark it complete.
- Read
metadata.md first.
- If
source/source_extract/ already exists with readable .tex files, use it directly.
- Otherwise, if
source/paper.pdf already exists, use PDF directly.
- If neither exists, run download scripts (single-paper scripts or batch script) first.
- Manually write
summary.md in the same paper directory, in the selected language.
Do not rely on rule-based auto summarization.
Do not rely on auto-extracted snippets as the primary writing basis.
Quality Requirement
- Every section should include paper-specific details that are traceable to full-text reading.
- Section 4/5/10 should reflect concrete method and evaluation details, not generic wording.
- If key details are unclear in the source, explicitly note uncertainty instead of guessing.
- Match the detail level shown in
references/summary-example-en.md and references/summary-example-zh.md.
- If your draft is clearly shorter or less specific than the examples, expand it before finishing.
Required Output
<paper_dir>/summary.md in fixed section format.
- Pay special attention to section
## 10. Brief Conclusion: write a 3-4 sentence mini-conclusion that covers contribution, method, evaluation setup, and results with paper-specific details.
- In section
## 1. Paper Snapshot, use exact keys: ArXiv ID, Title, Authors, Publish date, Primary category, Reading basis.
- Do not use key variants such as
Reading source, Author list, Published on, or lowercase key names.
See references/summary-format.md for exact section requirements.
Related Skills
This skill is a sub-skill of arxiv-summarizer-orchestrator.
Pipeline position:
- Step 1 (upstream):
arxiv-search-collector produces the selected paper directories and metadata.
- Step 2 (this skill):
arxiv-paper-processor downloads artifacts and writes one summary.md per paper.
- Step 3 (downstream):
arxiv-batch-reporter uses these per-paper summaries to generate the final collection report.
Use this skill together with Step 1 and Step 3 for full end-to-end execution.
1---2name: arxiv-paper-processor3description: Tool-only paper processing skill with a manual language parameter: supports batch artifact download for many papers or single-paper download, then the model manually reads source/PDF and writes summary.md in the selected language. Use when per-paper comprehension should be model-driven instead of script-generated.4---5
6# ArXiv Paper Processor
7
8Use this skill for per-paper manual summarization, with optional batch artifact download.
9
10- Single-paper mode: process one paper directory (e.g. `<run_dir>/<arxiv_id>/`).
11- Batch predownload mode: process many paper directories under one run dir before writing summaries.
12
13## Language Parameter
14
15- Use a workflow language parameter (for example `English` or `Chinese`) and apply it manually.
16- The per-paper `summary.md` must be written in the selected language.
17- If download scripts are called directly, pass `--language <LANG>` for traceability.
18
19## Core Principle
20
21Scripts only fetch artifacts. The model performs reading and writing.
22
23## Non-negotiable Constraint
24
25- Do not generate `summary.md` by script-based snippet extraction, regex harvesting, or template autofill.
26- Do not use Python/shell scripts to auto-compose section text from abstract/introduction fragments.
27- Scripts in this skill are only for artifact download (`source`/`pdf`) and trace logs.
28- The final `summary.md` must come from model-side reading and synthesis of the paper content.
29
30## Optional Batch Artifact Download (Many Papers)
31
32Use this first when Stage B has many papers:
33
34```bash
35python3 scripts/download_papers_batch.py \
36 --run-dir /path/to/run \
37 --artifact source_then_pdf \
38 --max-workers 3 \
39 --min-interval-sec 5 \
40 --language English
41```
42
43Key behavior:
44
45- Supports `--artifact source`, `--artifact pdf`, or `--artifact source_then_pdf` (default).
46- Supports concurrency (`--max-workers`) and safe throttling/retry (`--min-interval-sec`, retry args).
47- Uses run-local throttle state by default (`<run_dir>/.runtime/arxiv_download_state.json`) to reduce 429 risk.
48- Skips papers that already have usable `source/source_extract/*.tex` or existing `source/paper.pdf` (unless `--force`).
49- Resume-friendly: if a paper already has a completed `summary.md`, you can skip that paper's summary-writing step.
50- Writes batch log to `<run_dir>/download_batch_log.json` by default.
51
52## Step 1: Download Source (Preferred)
53
54```bash
55python3 scripts/download_arxiv_source.py \
56 --paper-dir /path/to/run/2602.00528 \
57 --language English
58```
59
60This writes:
61
62- `source/source_bundle.bin`
63- `source/source_extract/`
64- `source/download_source_log.json`
65
66If usable source already exists and `--force` is not set, the script reuses local artifacts.
67
68## Step 2: If Needed, Download PDF
69
70```bash
71python3 scripts/download_arxiv_pdf.py \
72 --paper-dir /path/to/run/2602.00528 \
73 --language English
74```
75
76This writes:
77
78- `source/paper.pdf`
79- `source/download_pdf_log.json`
80
81If PDF already exists and `--force` is not set, the script reuses local artifacts.
82
83## Step 3: Model Reads and Summarizes
84
851. If `summary.md` already exists and follows the required format, skip this paper and mark it complete.
862. Read `metadata.md` first.
873. If `source/source_extract/` already exists with readable `.tex` files, use it directly.
884. Otherwise, if `source/paper.pdf` already exists, use PDF directly.
895. If neither exists, run download scripts (single-paper scripts or batch script) first.
906. Manually write `summary.md` in the same paper directory, in the selected language.
91
92Do not rely on rule-based auto summarization.
93Do not rely on auto-extracted snippets as the primary writing basis.
94
95## Quality Requirement
96
97- Every section should include paper-specific details that are traceable to full-text reading.
98- Section 4/5/10 should reflect concrete method and evaluation details, not generic wording.
99- If key details are unclear in the source, explicitly note uncertainty instead of guessing.
100- Match the detail level shown in `references/summary-example-en.md` and `references/summary-example-zh.md`.
101- If your draft is clearly shorter or less specific than the examples, expand it before finishing.
102
103## Required Output
104
105- `<paper_dir>/summary.md` in fixed section format.
106- Pay special attention to section `## 10. Brief Conclusion`: write a 3-4 sentence mini-conclusion that covers contribution, method, evaluation setup, and results with paper-specific details.
107- In section `## 1. Paper Snapshot`, use exact keys: `ArXiv ID`, `Title`, `Authors`, `Publish date`, `Primary category`, `Reading basis`.
108- Do not use key variants such as `Reading source`, `Author list`, `Published on`, or lowercase key names.
109
110See `references/summary-format.md` for exact section requirements.
111
112## Related Skills
113
114This skill is a sub-skill of `arxiv-summarizer-orchestrator`.
115
116Pipeline position:
117
1181. Step 1 (upstream): `arxiv-search-collector` produces the selected paper directories and metadata.
1192. Step 2 (this skill): `arxiv-paper-processor` downloads artifacts and writes one `summary.md` per paper.
1203. Step 3 (downstream): `arxiv-batch-reporter` uses these per-paper summaries to generate the final collection report.
121
122Use this skill together with Step 1 and Step 3 for full end-to-end execution.