arXiv Search (metadata-first)
Triggers & routing
- Trigger: arXiv, arxiv, paper search, metadata retrieval, 文献检索, 论文检索, 拉取元数据, 离线导入.
- Use when: 需要一个初始论文集合(survey/snapshot 的 Stage C1),来源为 arXiv(在线检索或离线导入 export)。
Collect an initial paper set with enough metadata to support downstream ranking, taxonomy building, and citation generation.
When online, prefer rich arXiv metadata (categories, arxiv_id, pdf_url, published/updated, etc.). When offline, accept an export and convert it cleanly.
Load Order
Always read:
references/domain_pack_overview.md — how domain packs drive topic-specific behavior
Domain packs (loaded by topic match):
assets/domain_packs/llm_agents.json — pinned IDs, query rewrite rules for LLM agent topics
Script Boundary
Use scripts/run.py only for:
- arXiv API retrieval and XML parsing
- offline export conversion (CSV/JSON/JSONL normalization)
- metadata enrichment via
id_list backfill
Do not treat run.py as the place for:
- hardcoded topic detection or query rewriting (use domain packs)
- domain-specific pinned paper lists (externalize to
assets/domain_packs/)
Contract-driven behavior
- Domain-pack query rewriting is the default for broad discovery Workflows.
- A focused Workflow may set
quality_contract.retrieval_policy.domain_pack_query_mode: explicit; in that
mode, the query list in queries.md remains authoritative and the domain
pack must not replace its topic focus.
quality_contract.retrieval_policy.minimum_records turns a Workflow's raw
candidate-pool floor into a strict quality-gate check.
Input
queries.md (keywords, excludes, time window)
Outputs
papers/papers_raw.jsonl (JSONL; 1 paper per line)
- Each record includes at least:
title, authors, year, url, abstract
- When using the arXiv API online mode, records also include helpful metadata:
arxiv_id, pdf_url, categories, primary_category, published, updated, doi, journal_ref, comment
- Convenience index (optional but generated by the script):
Decision: online vs offline
- If you have network access: run arXiv API retrieval.
- If not: import an export the user provides (CSV/JSON/JSONL) and normalize fields.
- Hybrid: if you import offline but still have network later, you can enrich missing fields (abstract/authors/categories) via arXiv
id_list using --enrich-metadata or queries.md enrich_metadata: true.
Workflow (heuristic)
- Read
queries.md and expand into concrete query strings.
- Retrieve results (online) or import an export (offline).
- Normalize every record to include at least:
title, authors (array), year, url, abstract
- Keep the set broad at this stage; dedupe/ranking comes next.
- Apply time window and
max_results if specified.
Quality checklist
Side effects
- Allowed: create/overwrite
papers/papers_raw.jsonl; append notes to STATUS.md.
- Not allowed: write prose sections in
output/ before writing is approved.
Script
Quick Start
uv run python .codex/skills/arxiv-search/scripts/run.py --help
- Online:
uv run python .codex/skills/arxiv-search/scripts/run.py --workspace <workspace> --query "<query>" --max-results 200
- Offline import:
uv run python .codex/skills/arxiv-search/scripts/run.py --workspace <workspace> --input <export.csv|json|jsonl>
All Options
--query <q>: repeatable; multiple queries are unioned
--exclude <term>: repeatable; excludes applied after retrieval
--max-results <n>: cap total retrieved
--input <export.*>: offline mode (CSV/JSON/JSONL)
--enrich-metadata: best-effort enrich via arXiv id_list (needs network)
queries.md also supports: keywords, exclude, time window, max_results, enrich_metadata
Examples
- Online (multi-query + excludes):
uv run python .codex/skills/arxiv-search/scripts/run.py --workspace <workspace> --query "LLM agent" --query "tool use" --exclude "survey" --max-results 300
- Fetch a single paper by arXiv ID (direct
id_list fetch):
uv run python .codex/skills/arxiv-search/scripts/run.py --workspace <workspace> --query 2509.02547 --max-results 1
- Offline auto-detect (no flags):
- Place
papers/import.csv (or .json/.jsonl) under the workspace, then run: uv run python .codex/skills/arxiv-search/scripts/run.py --workspace <workspace>
- Offline import + time window (via
queries.md):
- Set
- time window: { from: 2022, to: 2025 } then run offline import normally
Troubleshooting
Common Issues
Issue: papers/papers_raw.jsonl is empty
Symptom:
- Script exits with “No results returned …” or output file is empty.
Causes:
- Network is blocked (online mode).
- Queries are too narrow or
queries.md is empty.
Solutions:
- Use offline import: place
papers/import.csv|json|jsonl in the workspace or pass --input.
- Broaden keywords and reduce excludes in
queries.md.
- Run with explicit
--query to sanity-check the parser.
Issue: Offline import records miss fields
Symptom:
- Downstream steps fail because records miss
authors/year/abstract/url.
Causes:
- Export columns don’t match expected fields; upstream export is incomplete.
Solutions:
- Ensure the export contains at least
title, authors, year, url, abstract.
- If you later have network, use
--enrich-metadata to backfill missing fields (best effort).
Recovery Checklist
1---2name: arxiv-search3description: Retrieve arXiv paper metadata with keyword queries or import an offline arXiv export, and save results as JSONL (`papers/papers_raw.jsonl`).4---56# arXiv Search (metadata-first)78## Triggers & routing910- **Trigger**: arXiv, arxiv, paper search, metadata retrieval, 文献检索, 论文检索, 拉取元数据, 离线导入.11- **Use when**: 需要一个初始论文集合(survey/snapshot 的 Stage C1),来源为 arXiv(在线检索或离线导入 export)。121314Collect an initial paper set with enough metadata to support downstream ranking, taxonomy building, and citation generation.1516When online, prefer rich arXiv metadata (categories, arxiv_id, pdf_url, published/updated, etc.). When offline, accept an export and convert it cleanly.1718## Load Order1920Always read:21- `references/domain_pack_overview.md` — how domain packs drive topic-specific behavior2223Domain packs (loaded by topic match):24- `assets/domain_packs/llm_agents.json` — pinned IDs, query rewrite rules for LLM agent topics2526## Script Boundary2728Use `scripts/run.py` only for:29- arXiv API retrieval and XML parsing30- offline export conversion (CSV/JSON/JSONL normalization)31- metadata enrichment via `id_list` backfill3233Do not treat `run.py` as the place for:34- hardcoded topic detection or query rewriting (use domain packs)35- domain-specific pinned paper lists (externalize to `assets/domain_packs/`)3637## Contract-driven behavior3839- Domain-pack query rewriting is the default for broad discovery Workflows.40- A focused Workflow may set41 `quality_contract.retrieval_policy.domain_pack_query_mode: explicit`; in that42 mode, the query list in `queries.md` remains authoritative and the domain43 pack must not replace its topic focus.44- `quality_contract.retrieval_policy.minimum_records` turns a Workflow's raw45 candidate-pool floor into a strict quality-gate check.4647## Input4849- `queries.md` (keywords, excludes, time window)5051## Outputs5253- `papers/papers_raw.jsonl` (JSONL; 1 paper per line)54 - Each record includes at least: `title`, `authors`, `year`, `url`, `abstract`55 - When using the arXiv API online mode, records also include helpful metadata: `arxiv_id`, `pdf_url`, `categories`, `primary_category`, `published`, `updated`, `doi`, `journal_ref`, `comment`56- Convenience index (optional but generated by the script):57 - `papers/papers_raw.csv`5859## Decision: online vs offline6061- If you have network access: run arXiv API retrieval.62- If not: import an export the user provides (CSV/JSON/JSONL) and normalize fields.63- Hybrid: if you import offline but still have network later, you can **enrich missing fields** (abstract/authors/categories) via arXiv `id_list` using `--enrich-metadata` or `queries.md` `enrich_metadata: true`.6465## Workflow (heuristic)66671. Read `queries.md` and expand into concrete query strings.682. Retrieve results (online) or import an export (offline).693. Normalize every record to include at least:70 - `title`, `authors` (array), `year`, `url`, `abstract`714. Keep the set broad at this stage; dedupe/ranking comes next.725. Apply time window and `max_results` if specified.7374## Quality checklist7576- [ ] `papers/papers_raw.jsonl` exists.77- [ ] Each line is valid JSON and contains `title`, `authors`, `year`, `url`.7879## Side effects8081- Allowed: create/overwrite `papers/papers_raw.jsonl`; append notes to `STATUS.md`.82- Not allowed: write prose sections in `output/` before writing is approved.8384## Script8586### Quick Start8788- `uv run python .codex/skills/arxiv-search/scripts/run.py --help`89- Online: `uv run python .codex/skills/arxiv-search/scripts/run.py --workspace <workspace> --query "<query>" --max-results 200`90- Offline import: `uv run python .codex/skills/arxiv-search/scripts/run.py --workspace <workspace> --input <export.csv|json|jsonl>`9192### All Options9394- `--query <q>`: repeatable; multiple queries are unioned95- `--exclude <term>`: repeatable; excludes applied after retrieval96- `--max-results <n>`: cap total retrieved97- `--input <export.*>`: offline mode (CSV/JSON/JSONL)98- `--enrich-metadata`: best-effort enrich via arXiv `id_list` (needs network)99- `queries.md` also supports: `keywords`, `exclude`, `time window`, `max_results`, `enrich_metadata`100101### Examples102103- Online (multi-query + excludes):104 - `uv run python .codex/skills/arxiv-search/scripts/run.py --workspace <workspace> --query "LLM agent" --query "tool use" --exclude "survey" --max-results 300`105- Fetch a single paper by arXiv ID (direct `id_list` fetch):106 - `uv run python .codex/skills/arxiv-search/scripts/run.py --workspace <workspace> --query 2509.02547 --max-results 1`107- Offline auto-detect (no flags):108 - Place `papers/import.csv` (or `.json/.jsonl`) under the workspace, then run: `uv run python .codex/skills/arxiv-search/scripts/run.py --workspace <workspace>`109- Offline import + time window (via `queries.md`):110 - Set `- time window: { from: 2022, to: 2025 }` then run offline import normally111112## Troubleshooting113114### Common Issues115116#### Issue: `papers/papers_raw.jsonl` is empty117118**Symptom**:119- Script exits with “No results returned …” or output file is empty.120121**Causes**:122- Network is blocked (online mode).123- Queries are too narrow or `queries.md` is empty.124125**Solutions**:126- Use offline import: place `papers/import.csv|json|jsonl` in the workspace or pass `--input`.127- Broaden keywords and reduce excludes in `queries.md`.128- Run with explicit `--query` to sanity-check the parser.129130#### Issue: Offline import records miss fields131132**Symptom**:133- Downstream steps fail because records miss `authors/year/abstract/url`.134135**Causes**:136- Export columns don’t match expected fields; upstream export is incomplete.137138**Solutions**:139- Ensure the export contains at least `title`, `authors`, `year`, `url`, `abstract`.140- If you later have network, use `--enrich-metadata` to backfill missing fields (best effort).141142### Recovery Checklist143144- [ ] Confirm `queries.md` has non-empty `keywords` (or pass `--query`).145- [ ] If offline: confirm workspace has `papers/import.*` and rerun.146- [ ] Spot-check 3–5 JSONL lines: valid JSON + required fields.