CrabPath
Pure graph core: zero required deps and no network calls. Caller provides callbacks.
Design Tenets
- No network calls in core
- No secret discovery (no dotfiles, keychain, or env probing)
- No subprocess provider wrappers
- Embedder identity in state metadata; dimension mismatches are errors
- One canonical state format (
state.json)
Quick Start
from crabpath import split_workspace, HashEmbedder, VectorIndex
graph, texts = split_workspace("./workspace")
embedder = HashEmbedder()
index = VectorIndex()
for nid, content in texts.items():
index.upsert(nid, embedder.embed(content))
Embeddings and LLM callbacks
- Default:
HashEmbedder (hash-v1, 1024-dim)
- Real: callback
embed_fn / embed_batch_fn (e.g., text-embedding-3-small)
- LLM routing: callback
llm_fn using gpt-5-mini (example)
Session Replay
replay_queries(graph, queries) can warm-start from historical turns.
CLI
--state is preferred:
crabpath query TEXT --state S [--top N] [--json]
crabpath query TEXT --state S --chat-id CID
crabpath doctor --state S
crabpath info --state S
crabpath init --workspace W --output O --embedder openai
crabpath query TEXT --state S --llm openai
crabpath inject --state S --type TEACHING [--type DIRECTIVE]
Real-time correction flow:
python3 query_brain.py --chat-id CHAT_ID
python3 learn_correction.py --chat-id CHAT_ID
Quick Reference
crabpath init/query/learn/inject/health/doctor/info
query_brain.py --chat-id and learn_correction.py for real-time correction pipelines
query_brain.py traversal limits: beam_width=8, max_hops=30, fire_threshold=0.01
- Hard traversal caps:
max_fired_nodes and max_context_chars (defaults None; query_brain.py defaults max_context_chars=20000)
examples/correction_flow/, examples/cold_start/, examples/openai_embedder/
API Reference
- Core lifecycle:
split_workspace
load_state
save_state
ManagedState
VectorIndex
- Traversal and learning:
traverse
TraversalConfig
TraversalConfig.beam_width, .max_hops, .fire_threshold, .max_fired_nodes, .max_context_chars, .reflex_threshold, .habitual_range, .inhibitory_threshold
TraversalResult
apply_outcome
- Runtime injection APIs:
inject_node
inject_correction
inject_batch
- Maintenance helpers:
suggest_connections, apply_connections
suggest_merges, apply_merge
measure_health, autotune, replay_queries
- Embedding utilities:
HashEmbedder
OpenAIEmbedder
default_embed
default_embed_batch
openai_llm_fn
- LLM routing callbacks:
- Graph primitives:
Node
Edge
Graph
split_workspace
generate_summaries
CLI Commands
crabpath init --workspace W --output O [--sessions S] [--embedder openai]
crabpath query TEXT --state S [--top N] [--json] [--chat-id CHAT_ID]
crabpath learn --state S --outcome N --fired-ids a,b,c [--json]
crabpath inject --state S --id NODE_ID --content TEXT [--type CORRECTION|TEACHING|DIRECTIVE] [--json] [--connect-min-sim 0.0]
crabpath inject --state S --id NODE_ID --content TEXT --type TEACHING
crabpath inject --state S --id NODE_ID --content TEXT --type DIRECTIVE
crabpath health --state S
crabpath doctor --state S
crabpath info --state S
crabpath replay --state S --sessions S
crabpath merge --state S [--llm openai]
crabpath connect --state S [--llm openai]
crabpath journal [--stats]
query_brain.py --chat-id CHAT_ID
learn_correction.py --chat-id CHAT_ID
Traversal defaults
beam_width=8
max_hops=30
fire_threshold=0.01
reflex_threshold=0.6
habitual_range=0.2-0.6
inhibitory_threshold=-0.01
max_fired_nodes (hard node-count cap, default None)
max_context_chars (hard context cap, default None; query_brain.py default is 20000)
Paper
https://jonathangu.com/crabpath/
1---2name: crabpath3description: Memory graph engine with caller-provided embed and LLM callbacks; core is pure, with real-time correction flow and optional OpenAI integration.4---5
6# CrabPath
7
8Pure graph core: zero required deps and no network calls. Caller provides callbacks.
9
10## Design Tenets
11
12- No network calls in core
13- No secret discovery (no dotfiles, keychain, or env probing)
14- No subprocess provider wrappers
15- Embedder identity in state metadata; dimension mismatches are errors
16- One canonical state format (`state.json`)
17
18## Quick Start
19
20```python
21from crabpath import split_workspace, HashEmbedder, VectorIndex
22
23graph, texts = split_workspace("./workspace")
24embedder = HashEmbedder()
25index = VectorIndex()
26for nid, content in texts.items():
27 index.upsert(nid, embedder.embed(content))
28```
29
30## Embeddings and LLM callbacks
31
32- Default: `HashEmbedder` (hash-v1, 1024-dim)
33- Real: callback `embed_fn` / `embed_batch_fn` (e.g., `text-embedding-3-small`)
34- LLM routing: callback `llm_fn` using `gpt-5-mini` (example)
35
36## Session Replay
37
38`replay_queries(graph, queries)` can warm-start from historical turns.
39
40## CLI
41
42`--state` is preferred:
43
44`crabpath query TEXT --state S [--top N] [--json]`
45`crabpath query TEXT --state S --chat-id CID`
46
47`crabpath doctor --state S`
48`crabpath info --state S`
49`crabpath init --workspace W --output O --embedder openai`
50`crabpath query TEXT --state S --llm openai`
51`crabpath inject --state S --type TEACHING [--type DIRECTIVE]`
52
53Real-time correction flow:
54`python3 query_brain.py --chat-id CHAT_ID`
55`python3 learn_correction.py --chat-id CHAT_ID`
56
57## Quick Reference
58- `crabpath init/query/learn/inject/health/doctor/info`
59- `query_brain.py --chat-id` and `learn_correction.py` for real-time correction pipelines
60- `query_brain.py` traversal limits: `beam_width=8`, `max_hops=30`, `fire_threshold=0.01`
61- Hard traversal caps: `max_fired_nodes` and `max_context_chars` (defaults `None`; `query_brain.py` defaults `max_context_chars=20000`)
62- `examples/correction_flow/`, `examples/cold_start/`, `examples/openai_embedder/`
63
64## API Reference
65
66- Core lifecycle:
67 - `split_workspace`
68 - `load_state`
69 - `save_state`
70 - `ManagedState`
71 - `VectorIndex`
72- Traversal and learning:
73 - `traverse`
74 - `TraversalConfig`
75 - `TraversalConfig.beam_width`, `.max_hops`, `.fire_threshold`, `.max_fired_nodes`, `.max_context_chars`, `.reflex_threshold`, `.habitual_range`, `.inhibitory_threshold`
76 - `TraversalResult`
77 - `apply_outcome`
78- Runtime injection APIs:
79 - `inject_node`
80 - `inject_correction`
81 - `inject_batch`
82- Maintenance helpers:
83 - `suggest_connections`, `apply_connections`
84 - `suggest_merges`, `apply_merge`
85 - `measure_health`, `autotune`, `replay_queries`
86- Embedding utilities:
87 - `HashEmbedder`
88 - `OpenAIEmbedder`
89 - `default_embed`
90 - `default_embed_batch`
91 - `openai_llm_fn`
92- LLM routing callbacks:
93 - `chat_completion`
94- Graph primitives:
95 - `Node`
96 - `Edge`
97 - `Graph`
98 - `split_workspace`
99 - `generate_summaries`
100
101## CLI Commands
102
103- `crabpath init --workspace W --output O [--sessions S] [--embedder openai]`
104- `crabpath query TEXT --state S [--top N] [--json] [--chat-id CHAT_ID]`
105- `crabpath learn --state S --outcome N --fired-ids a,b,c [--json]`
106- `crabpath inject --state S --id NODE_ID --content TEXT [--type CORRECTION|TEACHING|DIRECTIVE] [--json] [--connect-min-sim 0.0]`
107- `crabpath inject --state S --id NODE_ID --content TEXT --type TEACHING`
108- `crabpath inject --state S --id NODE_ID --content TEXT --type DIRECTIVE`
109- `crabpath health --state S`
110- `crabpath doctor --state S`
111- `crabpath info --state S`
112- `crabpath replay --state S --sessions S`
113- `crabpath merge --state S [--llm openai]`
114- `crabpath connect --state S [--llm openai]`
115- `crabpath journal [--stats]`
116- `query_brain.py --chat-id CHAT_ID`
117- `learn_correction.py --chat-id CHAT_ID`
118
119## Traversal defaults
120
121- `beam_width=8`
122- `max_hops=30`
123- `fire_threshold=0.01`
124- `reflex_threshold=0.6`
125- `habitual_range=0.2-0.6`
126- `inhibitory_threshold=-0.01`
127- `max_fired_nodes` (hard node-count cap, default `None`)
128- `max_context_chars` (hard context cap, default `None`; `query_brain.py` default is `20000`)
129
130## Paper
131
132https://jonathangu.com/crabpath/