# Arxiv Search

> Retrieve arXiv paper metadata with keyword queries or import an offline arXiv export, and save results as JSONL (`papers/papers_raw.jsonl`).

- Skill: `willoscar/arxiv-search` (Agent Skill, multi-file: 5 files)
- Install (CLI): `npx skillmds add willoscar/arxiv-search`
- Raw SKILL.md: https://api.skillmd.com/api/skills/willoscar/arxiv-search/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: WILLOSCAR (https://skillmd.com/u/willoscar)
- Updated: 2026-09-10
- Page: https://skillmd.com/skills/willoscar/arxiv-search

---


# arXiv Search (metadata-first)

## Triggers & routing

- **Trigger**: arXiv, arxiv, paper search, metadata retrieval, 文献检索, 论文检索, 拉取元数据, 离线导入.
- **Use when**: 需要一个初始论文集合（survey/snapshot 的 Stage C1），来源为 arXiv（在线检索或离线导入 export）。


Collect an initial paper set with enough metadata to support downstream ranking, taxonomy building, and citation generation.

When online, prefer rich arXiv metadata (categories, arxiv_id, pdf_url, published/updated, etc.). When offline, accept an export and convert it cleanly.

## Load Order

Always read:
- `references/domain_pack_overview.md` — how domain packs drive topic-specific behavior

Domain packs (loaded by topic match):
- `assets/domain_packs/llm_agents.json` — pinned IDs, query rewrite rules for LLM agent topics

## Script Boundary

Use `scripts/run.py` only for:
- arXiv API retrieval and XML parsing
- offline export conversion (CSV/JSON/JSONL normalization)
- metadata enrichment via `id_list` backfill

Do not treat `run.py` as the place for:
- hardcoded topic detection or query rewriting (use domain packs)
- domain-specific pinned paper lists (externalize to `assets/domain_packs/`)

## Contract-driven behavior

- Domain-pack query rewriting is the default for broad discovery Workflows.
- A focused Workflow may set
  `quality_contract.retrieval_policy.domain_pack_query_mode: explicit`; in that
  mode, the query list in `queries.md` remains authoritative and the domain
  pack must not replace its topic focus.
- `quality_contract.retrieval_policy.minimum_records` turns a Workflow's raw
  candidate-pool floor into a strict quality-gate check.

## Input

- `queries.md` (keywords, excludes, time window)

## Outputs

- `papers/papers_raw.jsonl` (JSONL; 1 paper per line)
  - Each record includes at least: `title`, `authors`, `year`, `url`, `abstract`
  - When using the arXiv API online mode, records also include helpful metadata: `arxiv_id`, `pdf_url`, `categories`, `primary_category`, `published`, `updated`, `doi`, `journal_ref`, `comment`
- Convenience index (optional but generated by the script):
  - `papers/papers_raw.csv`

## Decision: online vs offline

- If you have network access: run arXiv API retrieval.
- If not: import an export the user provides (CSV/JSON/JSONL) and normalize fields.
- Hybrid: if you import offline but still have network later, you can **enrich missing fields** (abstract/authors/categories) via arXiv `id_list` using `--enrich-metadata` or `queries.md` `enrich_metadata: true`.

## Workflow (heuristic)

1. Read `queries.md` and expand into concrete query strings.
2. Retrieve results (online) or import an export (offline).
3. Normalize every record to include at least:
   - `title`, `authors` (array), `year`, `url`, `abstract`
4. Keep the set broad at this stage; dedupe/ranking comes next.
5. Apply time window and `max_results` if specified.

## Quality checklist

- [ ] `papers/papers_raw.jsonl` exists.
- [ ] Each line is valid JSON and contains `title`, `authors`, `year`, `url`.

## Side effects

- Allowed: create/overwrite `papers/papers_raw.jsonl`; append notes to `STATUS.md`.
- Not allowed: write prose sections in `output/` before writing is approved.

## Script

### Quick Start

- `uv run python .codex/skills/arxiv-search/scripts/run.py --help`
- Online: `uv run python .codex/skills/arxiv-search/scripts/run.py --workspace <workspace> --query "<query>" --max-results 200`
- Offline import: `uv run python .codex/skills/arxiv-search/scripts/run.py --workspace <workspace> --input <export.csv|json|jsonl>`

### All Options

- `--query <q>`: repeatable; multiple queries are unioned
- `--exclude <term>`: repeatable; excludes applied after retrieval
- `--max-results <n>`: cap total retrieved
- `--input <export.*>`: offline mode (CSV/JSON/JSONL)
- `--enrich-metadata`: best-effort enrich via arXiv `id_list` (needs network)
- `queries.md` also supports: `keywords`, `exclude`, `time window`, `max_results`, `enrich_metadata`

### Examples

- Online (multi-query + excludes):
  - `uv run python .codex/skills/arxiv-search/scripts/run.py --workspace <workspace> --query "LLM agent" --query "tool use" --exclude "survey" --max-results 300`
- Fetch a single paper by arXiv ID (direct `id_list` fetch):
  - `uv run python .codex/skills/arxiv-search/scripts/run.py --workspace <workspace> --query 2509.02547 --max-results 1`
- Offline auto-detect (no flags):
  - Place `papers/import.csv` (or `.json/.jsonl`) under the workspace, then run: `uv run python .codex/skills/arxiv-search/scripts/run.py --workspace <workspace>`
- Offline import + time window (via `queries.md`):
  - Set `- time window: { from: 2022, to: 2025 }` then run offline import normally

## Troubleshooting

### Common Issues

#### Issue: `papers/papers_raw.jsonl` is empty

**Symptom**:
- Script exits with “No results returned …” or output file is empty.

**Causes**:
- Network is blocked (online mode).
- Queries are too narrow or `queries.md` is empty.

**Solutions**:
- Use offline import: place `papers/import.csv|json|jsonl` in the workspace or pass `--input`.
- Broaden keywords and reduce excludes in `queries.md`.
- Run with explicit `--query` to sanity-check the parser.

#### Issue: Offline import records miss fields

**Symptom**:
- Downstream steps fail because records miss `authors/year/abstract/url`.

**Causes**:
- Export columns don’t match expected fields; upstream export is incomplete.

**Solutions**:
- Ensure the export contains at least `title`, `authors`, `year`, `url`, `abstract`.
- If you later have network, use `--enrich-metadata` to backfill missing fields (best effort).

### Recovery Checklist

- [ ] Confirm `queries.md` has non-empty `keywords` (or pass `--query`).
- [ ] If offline: confirm workspace has `papers/import.*` and rerun.
- [ ] Spot-check 3–5 JSONL lines: valid JSON + required fields.

