/discover
Produce a ranked shortlist of paper candidates from one of three seed modes. Surface them to the user (or to the calling skill) with rationales. Never auto-ingest —
/discoveris a proposal stage,/ingestis the action stage.
Use these local references on demand:
references/seed-modes.md— when to pick anchor / topic / wiki mode and how to translate the user's phrasing into onereferences/ranking-signals.md— whattools/discover.pyscores on and why discovery does not share/init's survey preferencereferences/wiki-dedup.md— how candidates are filtered againstwiki/papers/and what to do with matches
Inputs
--anchor <id>(repeatable): one or more anchor paper IDs (arXiv IDs preferred; S2 paperIds also accepted). Drives the anchor mode — the primary use case, including the post-/ingest"what to read next" flow.--negative <id>(repeatable, optional): IDs to push recommendations away from. Only meaningful with--anchor.--topic "<str>": a topic / query string. Drives the topic mode — lighter alternative to/init's planner.--from-wiki: derive seeds automatically from the wiki's most recently modified papers. Drives the wiki mode.--limit N(optional, default 10): max shortlist size.
Exactly one of --anchor, --topic, --from-wiki must be given.
Outputs
.checkpoints/discover-{seed-slug}-{YYYY-MM-DD}.json— full shortlist payload, machine-readable; the seed slug is derived from the first anchor or the topic- a human-readable markdown summary printed to the user with rationale per candidate
wiki/log.md— one append line viatools/research_wiki.py log
/discover does not write anywhere else in wiki/ and does not touch raw/. Whether to actually pull a candidate into the wiki is the caller's decision (a follow-up /ingest).
Wiki Interaction
Reads
wiki/papers/*.md— frontmatterarxiv(or legacyarxiv_id) for dedup against already-ingested paperswiki/papers/*.mdmodification times — for--from-wikianchor selection
Writes
wiki/log.md— APPEND viatools/research_wiki.py log
Graph edges created
- none. Graph mutations belong to
/ingest, not/discover.
Workflow
Pre-condition: working directory contains wiki/, raw/, and tools/. Resolve the Python interpreter once and reuse it:
if [ -x .venv/bin/python ]; then
PYTHON_BIN=.venv/bin/python
elif [ -x .venv/Scripts/python.exe ]; then
PYTHON_BIN=.venv/Scripts/python.exe
else
PYTHON_BIN=python3
fi
export PYTHON_BIN
Step 1: Pick the seed mode
Translate the user's request into exactly one of from-anchors, from-topic, or from-wiki. The decision rule lives in references/seed-modes.md; the short version:
- the user named one or more specific papers, or this is a post-
/ingest--discoverfollow-up → anchors - the user gave a topic / direction / keywords → topic
- the user asked open-ended "what should I read next" with no anchor and no topic → wiki
If the user supplied negatives ("not these"), include them via --negative in anchor mode only.
Step 2: Run the discovery tool
"$PYTHON_BIN" tools/discover.py from-anchors \
--id <arxiv-id> [--id <arxiv-id>...] [--negative <id>...] \
--wiki-root wiki \
--limit 10 \
--output-checkpoint .checkpoints/ \
--markdown
Or for topic / wiki modes:
"$PYTHON_BIN" tools/discover.py from-topic "<query>" --wiki-root wiki --limit 10 --output-checkpoint .checkpoints/ --markdown
"$PYTHON_BIN" tools/discover.py from-wiki --wiki-root wiki --limit 10 --output-checkpoint .checkpoints/ --markdown
Anchor (and wiki) mode run three S2 channels per anchor by default — recommend + references + citations. This is what makes /discover meaningfully different from /daily-arxiv: references surface older canonical work the anchor built on, citations surface high-impact follow-ups. Pass --no-citation-expand only if API cost forces the narrower recommend-only path; the quality regression is sharp.
The tool handles candidate gathering, wiki dedup, ranking, and writes the checkpoint. Always pass --wiki-root wiki so already-ingested papers are filtered out — surfacing duplicates wastes the user's review time.
If S2 is unavailable in topic mode, the tool will continue with whatever sources responded; check the output and report degraded discovery to the user. If every channel fails, abort with a clear message rather than emitting an empty shortlist as if it were a real recommendation.
Step 3: Present the shortlist
Show the markdown output to the user. For each candidate, the user needs enough to decide whether to ingest:
- title and arXiv ID (or S2 paperId fallback)
- one-line rationale (already produced by the tool: anchor count, influential citations, year)
- TLDR if the tool surfaced one (topic-mode candidates often have it; anchor-mode usually does not — the recommendations endpoint does not return TLDRs)
Append a short "next step" hint:
To ingest a candidate: /ingest https://arxiv.org/abs/<arxiv-id>
Do not ingest anything yourself. The user picks.
Step 4: Log
"$PYTHON_BIN" tools/research_wiki.py log wiki "discover | mode=<anchors|topic|wiki> | seed=<short-desc> | shortlist=<N>"
Internal Callers
/discover is designed to be invoked both by users (manually) and by other skills (as a subroutine).
From /ingest --discover
When /ingest is invoked with the optional --discover flag (default off), it calls /discover after the final report, with the just-ingested paper's arXiv ID as the single anchor. The shortlist is appended to /ingest's report under a "Related papers you may want to ingest next" heading. /ingest never auto-ingests anything from this list.
From /init
/init does not call /discover. /init's planner (tools/init_discovery.py plan) has its own scoring that favors surveys, broad coverage, and seed anchors — appropriate for bootstrapping a wiki. /discover's ranking is intentionally different (no survey preference; weights anchor similarity and influential citations) and would dilute /init's shortlist if substituted in. Keep them separate.
Constraints
- Never auto-ingest:
/discoverreturns a shortlist and stops. Even when called by/ingest --discover, the caller surfaces results and the user decides what to ingest. - No writes to
wiki/other thanlog.md: paper pages, concepts, claims, graph edges all belong to/ingest. - No writes to
raw/:/discoverdoes not download papers. The user runs/ingest <arxiv-url>afterwards if they want a candidate. - Always dedupe against the wiki: pass
--wiki-root wikiso the shortlist contains only papers not yet in the wiki. Surfacing duplicates is the most common low-quality failure mode. - Ranking is discovery-specific: do not import or duplicate
tools/init_discovery.py's scoring helpers. The two skills have different objectives —/initwants broad foundational coverage;/discoverwants relevant next reads. Seereferences/ranking-signals.md. - Three-channel anchor gather: by default, anchor mode pulls from S2
recommend+references+citationsper anchor. Removing the citation channels (via--no-citation-expand) collapses the result into a recency-biased semantic cluster that overlaps heavily with/daily-arxiv. Keep all three on unless API cost is a hard constraint. Seereferences/ranking-signals.md. - Some S2 endpoints have a flatter field set:
/citations,/references, and/recommendations/*reject nested selectors — noauthors.hIndex, notldr./paper/{id}and/paper/searchdo accept them, so topic-mode candidates carry full enrichment; anchor-mode candidates that entered only via citations/references/recommend do not. That is a real API constraint, not a bug. - Rate limits apply: each anchor in anchor mode costs up to three S2 calls (recommend + references + citations). Default per-anchor limit is 50 for recs and 30 each for references/citations. Multi-anchor runs multiply accordingly; with an API key (1 req/sec) a 3-anchor run takes ~10 seconds.
Error Handling
- All seed channels fail: report the failure, write no shortlist, and do not log a successful run.
- S2 unavailable, DeepXiv available (topic mode): continue with DeepXiv only; note the degradation in the report.
- S2 returns zero recommendations for an anchor: keep going with the remaining anchors; if all anchors return zero, treat as total failure.
--from-wikifinds no anchorable papers (wiki/papers/empty or all missingarxiv_id): tell the user the wiki is too sparse for wiki-mode discovery and suggest topic mode.- Anchor ID is malformed or unknown: S2 will return 404; surface the bad ID in the report and continue with any remaining anchors.
Dependencies
Tools (via Bash)
"$PYTHON_BIN" tools/discover.py from-anchors --id <id> [--id <id>...] [--negative <id>...] --wiki-root wiki --limit <N> --output-checkpoint .checkpoints/ --markdown"$PYTHON_BIN" tools/discover.py from-topic "<query>" --wiki-root wiki --limit <N> --output-checkpoint .checkpoints/ --markdown"$PYTHON_BIN" tools/discover.py from-wiki --wiki-root wiki --limit <N> --output-checkpoint .checkpoints/ --markdown"$PYTHON_BIN" tools/research_wiki.py log wiki "<message>"
Skills
/ingest— caller via--discoverflag; also the action the user takes on a chosen candidate/init— independent planner; does not call/discover
External APIs
- Semantic Scholar — recommendations (
/recommendations/v1/papers/forpaper/{id},POST /recommendations/v1/papers/), search, paper detail (viatools/fetch_s2.py) - DeepXiv — search fallback in topic mode (via
tools/fetch_deepxiv.py, optional; graceful fallback when unavailable)