# Vault Ingest

> Ingest sources into the Obsidian wiki vault. Reads a source, extracts entities and concepts, creates or updates wiki pages, cross-references, and logs the operation. Supports files, URLs, and batch mode. Triggers on: ingest, process this source, add this to the wiki, read and file this, batch ingest, ingest all of these, ingest this url.

- Skill: `danmestas/vault-ingest` (Agent Skill)
- Install (CLI): `npx skillmds@latest add danmestas/vault-ingest`
- Raw SKILL.md: https://api.skillmd.com/api/skills/danmestas/vault-ingest/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: danmestas (https://skillmd.com/u/danmestas)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/danmestas/vault-ingest

---


# wiki-ingest: Source Ingestion

Read the source. Write the wiki. Cross-reference everything. A single source typically touches 8-15 wiki pages.

**Syntax standard**: Write all Obsidian Markdown using proper Obsidian Flavored Markdown. Wikilinks as `[[Note Name]]`, callouts as `> [!type] Title`, embeds as `![[file]]`, properties as YAML frontmatter. If the kepano/obsidian-skills plugin is installed, prefer its canonical Obsidian Markdown reference for syntax guidance. Otherwise, follow the guidance in this skill.

---

## Delta Tracking

Before ingesting any file, check `raw/.manifest.json` to avoid re-processing unchanged sources.

```bash
# Check if manifest exists
[ -f raw/.manifest.json ] && echo "exists" || echo "no manifest yet"
```

**Manifest format** (create if missing):
```json
{
  "sources": {
    "raw/articles/article-slug-2026-04-08.md": {
      "hash": "abc123",
      "ingested_at": "2026-04-08",
      "pages_created": ["wiki/sources/article-slug.md", "wiki/entities/Person.md"],
      "pages_updated": ["wiki/index.md"]
    }
  }
}
```

**Before ingesting a file:**
1. Compute a hash: `md5sum [file] | cut -d' ' -f1` (or `sha256sum` on Linux).
2. Check if the path exists in `.manifest.json` with the same hash.
3. If hash matches, skip. Report: "Already ingested (unchanged). Use `force` to re-ingest."
4. If missing or hash differs, proceed with ingest.

**After ingesting a file:**
1. Record `{hash, ingested_at, pages_created, pages_updated}` in `.manifest.json`.
2. Write the updated manifest back.

Skip delta checking if the user says "force ingest" or "re-ingest".

---

## URL Ingestion

Trigger: user passes a URL starting with `https://`.

Steps:

1. **Fetch** the page using WebFetch.
2. **Clean** (optional): if `defuddle` is available (`which defuddle 2>/dev/null`), run `defuddle [url]` to strip ads, nav, and clutter. Typically saves 40-60% tokens. Fall back to raw WebFetch output if not installed.
3. **Derive slug** from the URL path (last segment, lowercased, spaces→hyphens, strip query strings).
4. **Save** to `raw/articles/[slug]-[YYYY-MM-DD].md` with a frontmatter header:
   ```markdown
   ---
   source_url: [url]
   fetched: [YYYY-MM-DD]
   ---
   ```
5. Proceed with **Single Source Ingest** starting at step 2 (file is now in `raw/`).

---

## Image / Vision Ingestion

Trigger: user passes an image file path (`.png`, `.jpg`, `.jpeg`, `.gif`, `.webp`, `.svg`, `.avif`).

Steps:

1. **Read** the image file using the Read tool. Claude can process images natively.
2. **Describe** the image contents: extract all text (OCR), identify key concepts, entities, diagrams, and data visible in the image.
3. **Save** the description to `raw/images/[slug]-[YYYY-MM-DD].md`:
   ```markdown
   ---
   source_type: image
   original_file: [original path]
   fetched: YYYY-MM-DD
   ---
   # Image: [slug]

   [Full description of image contents, transcribed text, entities visible, etc.]
   ```
4. Copy the image to `_attachments/images/[slug].[ext]` if it's not already in the vault.
5. Proceed with **Single Source Ingest** on the saved description file.

Use cases: whiteboard photos, screenshots, diagrams, infographics, document scans.

---

## Single Source Ingest

Trigger: user drops a file into `raw/` or pastes content.

Steps:

1. **Read** the source completely. Do not skim.
2. **Discuss** key takeaways with the user. Ask: "What should I emphasize? How granular?" Skip this if the user says "just ingest it."
3. **Create** source summary in `wiki/sources/`. Use the source frontmatter schema from `references/frontmatter.md`.
4. **Create or update** entity pages for every person, org, product, and repo mentioned. One page per entity.
5. **Create or update** concept pages for significant ideas and frameworks.
6. **Update** relevant domain page(s) and their `_index.md` sub-indexes.
7. **Update** `wiki/overview.md` if the big picture changed.
8. **Update** `wiki/index.md`. Add entries for all new pages.
9. **Update** `wiki/hot.md` with this ingest's context.
10. **Append** to `wiki/log.md` (new entries at the TOP):
    ```markdown
    ## [YYYY-MM-DD] ingest | Source Title
    - Source: `raw/articles/filename.md`
    - Summary: [[Source Title]]
    - Pages created: [[Page 1]], [[Page 2]]
    - Pages updated: [[Page 3]], [[Page 4]]
    - Key insight: One sentence on what is new.
    ```
11. **Check for contradictions.** If new info conflicts with existing pages, add `> [!contradiction]` callouts on both pages.
12. **Move the original out of the inbox.** After the wiki pages are written, move the source file from `raw/` to `_sources/<category>/` (choose the category subfolder by kind — `articles/`, `research/`, `specs/`, `notes/`, ...), or to `_archive/` if it is large (>100 KB) or low-value. `raw/` is an inbox and should end empty once everything is processed.
13. **Sync the qkb index.** Run `qkb update` so BM25, graph link extraction, and the question-cache used by `vault-query-qkb` pick up everything you just wrote. See **Index Sync** below.

---

## Batch Ingest

Trigger: user drops multiple files or says "ingest all of these."

Steps:

1. List all files to process. Confirm with user before starting.
2. Process each source following the single ingest flow. Defer cross-referencing between sources until step 3.
3. After all sources: do a cross-reference pass. Look for connections between the newly ingested sources.
4. Update index, hot cache, and log once at the end (not per-source).
5. Report: "Processed N sources. Created X pages, updated Y pages. Here are the key connections I found."
6. **Sync the qkb index once at the end** (not per-source). See **Index Sync** below.

Batch ingest is less interactive. For 30+ sources, expect significant processing time. Check in with the user after every 10 sources.

---

## Context Window Discipline

Token budget matters. Follow these rules during ingest:

- Read `wiki/hot.md` first. If it contains the relevant context, don't re-read full pages.
- Read `wiki/index.md` to find existing pages before creating new ones.
- Read only 3-5 existing pages per ingest. If you need 10+, you are reading too broadly.
- Use PATCH for surgical edits. Never re-read an entire file just to update one field.
- Keep wiki pages short. 100-300 lines max. If a page grows beyond 300 lines, split it.
- Use search (`/search/simple/`) to find specific content without reading full pages.

---

## Contradictions

> [!note] Custom callout dependency
> The `[!contradiction]` callout type used below is a **custom callout** defined in `.obsidian/snippets/vault-colors.css` (auto-installed by `/wiki` scaffold). It renders with reddish-brown styling and an alert-triangle icon when the snippet is enabled. If the snippet is missing, Obsidian falls back to default callout styling, so the page still works without the visual flourish. See [[skills/wiki/references/css-snippets.md]] for the four custom callouts (`contradiction`, `gap`, `key-insight`, `stale`).

When new info contradicts an existing wiki page:

On the existing page, add:
```markdown
> [!contradiction] Conflict with [[New Source]]
> [[Existing Page]] claims X. [[New Source]] says Y.
> Needs resolution. Check dates, context, and primary sources.
```

On the new source summary, reference it:
```markdown
> [!contradiction] Contradicts [[Existing Page]]
> This source says Y, but existing wiki says X. See [[Existing Page]] for details.
```

Do not silently overwrite old claims. Flag and let the user decide.

---

## Naming Conventions

Obsidian resolves wikilinks by matching the **exact filename** (without extension). Mismatched names create broken links.

| Page Type | Filename Convention | Example | Wikilink |
|-----------|-------------------|---------|----------|
| **Entity** | Title Case with spaces | `Universal Weather.md` | `[[Universal Weather]]` |
| **Concept** | Title Case with spaces | `Missing Middle.md` | `[[Missing Middle]]` |
| **Domain** | Title Case with spaces | `Competitive Landscape.md` | `[[Competitive Landscape]]` |
| **Comparison** | Title Case with spaces | `Us vs Universal Weather.md` | `[[Us vs Universal Weather]]` |
| **Question** | Title Case with spaces | `Research NOTAM Reform.md` | `[[Research NOTAM Reform]]` |
| **Source** | kebab-case (slug) | `competitor-research-deep-dive.md` | `[[competitor-research-deep-dive\|Display Name]]` |
| **Meta** | lowercase | `index.md`, `log.md`, `hot.md` | `[[index]]`, `[[log]]`, `[[hot]]` |

**Why the split?** Entities, concepts, domains, comparisons, and questions are referenced frequently with natural-language wikilinks (`[[Missing Middle]]`). The filename must match exactly. Source summaries are referenced less frequently and use pipe syntax for display names (`[[some-slug|Readable Title]]`), so kebab-case slugs are fine.

**The rule:** If the wikilink text is the page name people type, the filename must match it exactly. If the wikilink always uses pipe syntax, kebab-case is acceptable.

---

## Index Sync

After every ingest (single or batch), the qkb retrieval index needs to know about the new/changed pages. Run this once at the end of the ingest:

```bash
qkb update
```

What this does:
- **Walks the vault** and diffs file content hashes against the index — only changed/new files get re-processed.
- **Rebuilds BM25** entries for changed docs so lexical search picks them up.
- **Auto-runs `graph link`** internally — extracts `[[wikilinks]]`, `![[embeds]]`, and frontmatter `type`s into the typed graph (LINKS_TO / EMBEDS / REFERENCES edges). This is what the `--graph` retrieval mode in `vault-query-qkb` relies on.
- **Idempotent**: re-running with no vault changes is a near-instant no-op.

Latency: typically 1–5 seconds for a handful of new pages, 30–60 seconds for a batch of 50+ new files. Print the line to the user so they can confirm freshness.

### Vector embeddings (the slow step) — NOT automatic

`qkb update` only refreshes lexical + graph state. New vector embeddings for the new docs are a separate, slower step (uses local LLM, 5–15 minutes for 100+ docs):

```bash
qkb embed
```

Do NOT run `embed` automatically inside the ingest skill. Two reasons:

1. **Lexical + graph alone is usually enough.** The reranker in `qkb query` scores against fresh chunk text from the FTS5 index regardless of vector freshness, so newly-ingested docs still surface correctly on most queries. They're slightly weaker on pure-semantic matches until embeddings catch up.
2. **It's heavy.** A user doing 20 small ingests across a session shouldn't trigger 20 model-load + embed cycles.

Instead, after the ingest report, **suggest** to the user:

> "Lexical + graph synced. Run `qkb embed` when you want the new pages' vectors caught up (5-15 min, optional)."

### When `qkb update` fails

Don't block the ingest report on this. If `qkb update` returns non-zero:

- Log the error briefly: "qkb index sync failed: <one-line>. Ingest itself succeeded; retry with `qkb update` once resolved."
- Continue with the rest of the ingest output.

Common causes: qkb not on PATH (user hasn't installed `@agent-ops/qkb` globally), GraphQLite extension missing, or a transient SQLite lock from a running MCP daemon.

---

## What Not to Do

- `raw/` is an inbox: after processing, move the original to `_sources/<category>/` (or `_archive/`), leaving `raw/` empty. Do not treat `raw/` as permanent storage.
- Do not modify the archived originals in `_sources/` or `_archive/`. Those are the immutable source documents.
- Do not create duplicate pages. Always check the index and search before creating.
- Do not skip the log entry. Every ingest must be recorded.
- Do not skip the hot cache update. It is what keeps future sessions fast.
- Do not skip the qkb index sync. Future `vault-query-qkb` calls will miss what you just wrote.
- Do not use kebab-case filenames for entities, concepts, domains, comparisons, or questions. These must match wikilink text exactly (Title Case with spaces).
- Do not run `qkb embed` automatically inside an ingest. Suggest it; let the user trigger.

