NuggetIndex: time-aware, conflict-aware fact governance
You apply the NuggetIndex method (Zerhoudi et al., SIGIR '26): represent knowledge as atomic facts ("nuggets") with explicit temporal validity, detect when sources disagree, and surface conflicts instead of silently picking a winner.
Core model
A nugget is one atomic fact:
- Triple:
(subject, predicate, object)+ the original supporting text span - Validity: interval
[start, end)—end = nullmeans "still true as far as we know" - Lifecycle:
active(current) |deprecated(obsolete) |contested(sources disagree) - Provenance: which source(s) assert it
Three failure modes this prevents:
- Temporal staleness — answering "Who is Apple's CEO?" with Steve Jobs because an old
passage ranked highest. Every fact carries
[start, end); check it against query time. - Silent conflicts — two sources claim different values for the same single-valued
attribute; naive RAG picks whichever embeds better. Mark both
contestedand tell the user. - Rename drift — Twitter→X, Facebook→Meta. Track renames as facts (
renamedTo) so both surface forms resolve to one entity.
Decision ladder — cheapest rung that fits
Rung 1: In-context audit (most common — no files, no install)
The user shows you retrieved passages, an answer that looks outdated, or asks "is my RAG returning stale facts?":
- Establish the query time (default: today; the user may ask about a past time).
- Extract the fact triples relevant to the query from each passage, with any temporal cues — follow references/extraction.md.
- Infer each fact's validity interval from its temporal cues (rules in extraction.md).
- Group facts by key
(subject, predicate). Apply the conflict rules in references/conflict-rules.md: single-valued predicate + overlapping validity + different objects ⇒ contested. - Report per passage:
consistent/stale(validity ended before query time) /contested(disagrees with another passage), with the evidence spans.
Never resolve a contest yourself when evidence is symmetric — show both claims, their sources, and their intervals, and let the user adjudicate.
Rung 2: Project fact ledger (ongoing work, small working set)
The user wants governed facts across a session or project (docs, research notes, a small
knowledge base): maintain nuggets.jsonl — one nugget per line in the exact schema of
references/ledger-format.md. The format is identical to the
pip package's Nugget JSON, so the ledger imports into a real index later with no migration.
On every append: extract → canonicalize entity names (resolve known renames/aliases to one
canonical form) → infer validity → dedup (same subject|predicate|object|start|scope = same
fact; merge provenance) → run conflict detection against existing ledger entries and update
lifecycle states. Report any new contested keys to the user immediately.
Rung 3: Escalate to the pip package (scale, persistence, serving)
Switch when any of these hold: corpus beyond ~50 documents, need for persistent hybrid BM25+dense retrieval, an HTTP sidecar for a production RAG stack, LangChain/LlamaIndex/ Haystack integration, or reproducible evaluation.
pip install nuggetindex # extras: [openai] [anthropic] [dense] [serve] [all]
nuggetindex build ./corpus --db index.db
nuggetindex query "CEO Apple" --db index.db --time 2010-06-01
Full CLI/API surface, extras, and a hand-ledger import recipe:
references/package.md. LLM-backed extraction needs an API key
(OPENAI_API_KEY / ANTHROPIC_API_KEY / GOOGLE_API_KEY, or local Ollama); the
rule-based extractor works with none. If no key is available and rule-based extraction is
too weak for the corpus, say so — do not fabricate extraction results.
Rules that always apply
- Timezone-aware UTC datetimes everywhere. Intervals are half-open
[start, end). - Query-time semantics: a fact answers a query at time
tiff its validity containstand its status isactiveorcontested— neverdeprecated. - Surface, don't hide: contested facts are presented as disputes with provenance, e.g.
Apple CEO: Tim Cook (2014–present) [DISPUTED: source B claims John Sculley (overlapping)]. - Be literal in extraction: no invented facts, no paraphrased evidence spans, omit facts you are not confident about.
- Unknown predicates default to multi-valued (conservative — avoids false conflicts). Only the single-valued predicates listed in conflict-rules.md trigger contests by default.