Developer Guide: WilmerAI Memory System
This guide provides a deep dive into the architecture and implementation of the conversational memory features within
WilmerAI. It has been updated to reflect the powerful workflow-based memory generation system, the specifics of the
vector search implementation, and the complete data schemas as verified by the current codebase.
1. Core Concepts & Architecture
WilmerAI's memory system is a sophisticated, multi-layered feature set implemented through specialized nodes within the
workflow engine. This design allows for different types of memory operations—creation, retrieval, and summarization—to
be strategically placed within workflows.
Key Architectural Principles:
Activation via discussionId: All persistent, stateful memory features are tied to a discussionId. Its presence
activates persistent storage mechanisms; its absence causes memory nodes to fall back to stateless, in-memory
operations using the current chat history.
Separation of Concerns: Creators vs. Retrievers: The system is fundamentally split into two categories of
operations, handled by distinct services:
- Memory Creation (Write): These are computationally expensive processes that analyze conversation history,
generate new summarized memories, and write them to persistent storage. This is the exclusive responsibility of
the
SlowButQualityRAGTool. This tool can generate memories either by executing a full sub-workflow for
complex logic or by making a direct LLM call.
- Memory Retrieval (Read): These are inexpensive, fast operations that read from storage or the current chat
history to provide context for an LLM. This is the primary responsibility of the
MemoryService.
Node-Based Implementation: Each memory operation is defined as a specific node_type in a workflow's JSON
configuration. This gives developers explicit control over when to perform expensive write operations versus cheap
read operations.
Centralized Routing and Registration:
- The
MemoryNodeHandler acts as a central router, directing requests for different memory node_types to the
appropriate service (MemoryService or SlowButQualityRAGTool).
- The
WorkflowManager acts as the system registrar where new memory node_types must be mapped to the
MemoryNodeHandler.
Persistent Storage: When a discussionId is active, the system maintains state in a user-specific directory using
a set of discussion-specific files:
- Memory File (
<id>_memories.jsonl): Stores discrete, summarized chunks of the conversation for file-based
memory. Each chunk is saved with a hash of the last message it's based on, creating a traceable, append-only
ledger.
- Chat Summary File (
<id>_summary.jsonl): Stores a single, continuously updated "rolling summary" of the
entire conversation. It's linked via a hash to the last memory chunk from the memory file that it incorporated.
- Vector Memory Database (
<id>_vector_memory.db): A discussion-specific SQLite database created on-demand
for each discussionId. It uses the FTS5 extension for powerful, weighted, full-text search across two main
tables:
memories table: Stores the ground-truth data. For vector memories, the memory_text column holds the *
LLM-generated summary*, not the raw conversation chunk. It also stores the full metadata_json.
memories_fts table: A virtual table that indexes the metadata for fast searching. The indexed columns
are title, summary, entities, key_phrases, and the original memory_text. Search relevance is
determined by the bm25 ranking function.
- Recency Scoring: The database connection is initialized with a custom SQL function,
recency_score, which
can calculate a time-decay boost for memories. While the default search query uses only bm25 ranking, this
function is available for developers to implement time-sensitive search ranking logic.
- Vector Memory Tracker (
vector_memory_tracker table): Located inside the <id>_vector_memory.db, this table
stores the hash of the last message processed for vector memory creation. This crucial feature prevents the
system from re-processing the same conversation history on subsequent runs.
2. Anatomy of Memory Node Types
Understanding the creator/retriever pattern clarifies the role of each node type.
Memory Creation & Persistence (Write)
These nodes perform the "heavy lifting" of creating and saving memories. They are powered by **SlowButQualityRAGTool
**.
QualityMemory: This is the primary memory creator node. It's designed to be run periodically in a workflow
to keep persistent memory up-to-date.
- Process Flow:
- The node triggers
SlowButQualityRAGTool.handle_discussion_id_flow.
- The tool checks if new memories are needed by comparing the current conversation history against the last
processed point.
- (Correction Start) It checks the
useVectorForQualityMemory flag from the config to determine its
behavior: if true, it generates vector memories; if false, it generates file-based memories. This
flag controls the behavior of the QualityMemory node specifically. Other nodes (like RecentMemory) can
trigger the creation of file-based memories independently of this flag's value. (Correction End)
- For the chosen memory type, the tool generates memories using one of two methods, based on its configuration:
- A) Workflow-Based Generation (Recommended) ✨: If the config specifies a
fileMemoryWorkflowName or
vectorMemoryWorkflowName, the tool executes that sub-workflow.
- File-Based Workflow: Passes the raw text chunk, recent memories, full memories, and the current
chat summary as
scoped_inputs. The workflow's final output **must be a single summarized text block
**.
- Vector-Based Workflow: Passes only the raw text chunk as a
scoped_input. The workflow's final
output must be a JSON string representing a single object or an array of objects.
- B) Direct LLM Call (Legacy): If a workflow name is not provided, the system falls back to a direct LLM
call using prompts defined in the config.
Memory Retrieval (Read)
These nodes perform fast, inexpensive "read" operations and are powered by MemoryService.
RecentMemory & RecentMemorySummarizerTool: The primary file-based memory retriever.
- With
discussionId (Stateful): Calls MemoryService.get_recent_memories, which reads the last N memory
chunks from <id>_memories.jsonl.
- Without
discussionId (Stateless): Falls back to an in-memory operation, grabbing the last N turns from the
current messages list.
VectorMemorySearch: The primary RAG search node.
- Purpose: Performs a highly relevant, keyword-based search against the discussion-specific vector memory
database.
- Process Flow: This node takes a string of keywords. The keywords must be separated by semicolons (
;).
The MemoryService calls vector_db_utils.search_memories_by_keyword.
- Keyword Handling: Internally, this process is robust:
- Sanitization: Each keyword is passed through the
_sanitize_fts5_term function, which wraps it in double
quotes (") to handle multi-word phrases and prevent FTS5 syntax errors.
- Query Construction: The sanitized terms are combined into a
MATCH query using OR logic.
- Limits: The system truncates the keyword list to a maximum of
60 (MAX_KEYWORDS_FOR_SEARCH) to avoid
exceeding SQLite's expression depth limit.
- Ranking: The final results are ranked by relevance using the
bm25 algorithm.
FullChatSummary: Retrieves the holistic, rolling summary of the conversation from <id>_summary.jsonl.
3. Critical Files for Development
To modify or extend the memory system, you will primarily work with these five files:
Middleware/workflows/tools/slow_but_quality_rag_tool.py: The heart of memory creation. Modify this file to
change how memories are generated, including the logic for choosing between vector/file and workflow/LLM-call
methods.
Middleware/services/memory_service.py: The core of memory retrieval. Modify this file to change how
memories are read from .jsonl files or retrieved from the vector database.
Middleware/utilities/vector_db_utils.py: The database abstraction layer. This contains all logic for
interacting with the discussion-specific SQLite databases, including table creation, data insertion, FTS5 search
query construction, and term sanitization.
Middleware/workflows/handlers/impl/memory_node_handler.py: The central router. You must update this file's
handle method to route any new node_type you create to the correct service or tool.
Middleware/workflows/managers/workflow_manager.py: The system registrar. You must register your new
node_type in the node_handlers dictionary within this manager's constructor. This manager is also responsible for
executing the memory generation sub-workflows.
4. Memory Generation Triggering Logic
This section details the internal logic that determines when new file-based memories are generated. Understanding this
is essential for debugging memory timing issues.
Key Configuration Fields
The triggering logic is governed by four fields in the discussion ID workflow configuration file:
chunkEstimatedTokenSize (default: 1000): The estimated token count threshold. When the accumulated new messages
reach or exceed this many tokens, memory generation triggers.
maxMessagesBetweenChunks (default: 5): The message count threshold. When this many new messages have
accumulated, memory generation triggers. This threshold only applies when the memory file exists on disk
(standard mode). It is disabled when the file does not exist (consolidation mode).
maxResponseSizeInTokens (default: 400): The maximum number of tokens in the LLM's summarization response.
lookbackStartTurn (default: 3): The number of most recent conversation turns to exclude from memory processing.
This prevents recent, potentially incomplete exchanges (such as frontend command messages) from being included in
memories.
Triggering Modes
The system operates in two modes depending on whether the memory file exists on disk:
Standard Mode (Memory File Exists)
Both thresholds are active. Memory generation triggers when either threshold is reached first:
estimated_tokens >= chunkEstimatedTokenSize, OR
message_count >= maxMessagesBetweenChunks
This "whichever comes first" logic allows the system to generate memories based on either volume (tokens) or frequency
(messages), depending on which condition is met first.
The memory file does not need to contain any chunks for standard mode to apply — it just needs to exist on disk. In a
new conversation, the file is created (as an empty []) on the first check (around message 3, after the early return
for conversations with fewer than 3 messages). From that point on, standard mode applies and both thresholds are active.
Consolidation Mode (Memory File Does Not Exist)
Only the token threshold applies. The message count threshold is disabled. This mode activates only when the memory file
has been deleted from disk — for example, when a user wants to regenerate memories with a larger chunk size to produce
fewer, larger chunks.
Implementation Details
The triggering logic lives in SlowButQualityRAGTool.handle_discussion_id_flow(), within the file-based memory path
(the else branch after the use_vector_memory_from_config check).
- File existence check: Before reading, the system checks
os.path.exists(filepath) to determine whether the
memory file exists on disk. This check must happen before read_chunks_with_hashes(), which creates the file as an
empty [] if it does not exist (see "File Creation Side-Effect" below).
- New message detection: The method uses hash-based tracking (
find_last_matching_hash_message) to identify
messages that have been added since the last memory generation run. Only unprocessed messages are considered. If the
file has no chunks (new chat or freshly created file), all messages minus the lookback are treated as new.
- Token estimation: The content of all new messages is joined and passed through
rough_estimate_token_length() to
get an approximate token count. This function intentionally overestimates by using the higher of a word-based estimate
(1.35 tokens/word) and a character-based estimate (3.5 chars/token), then applying a safety_margin multiplier
(default 1.10). See text_utils.py for details.
- Trigger evaluation: The two boolean conditions are evaluated.
trigger_by_messages is gated on file_exists
(from step 1). Diagnostic logging reports the exact counts, thresholds, and which trigger(s) fired (or why none did).
File Creation Side-Effect
The read_chunks_with_hashes() function calls ensure_json_file_exists(), which creates the memory file as an empty
JSON array ([]) if it does not exist on disk. This is why the file_exists check must happen before that call —
otherwise the file would always appear to exist. The side-effect is intentional: it means that on the first memory
check of a new conversation (around message 3), the file is created, and from message 4 onward the system operates in
standard mode with both thresholds active.
5. How to Add a New Memory Feature
The system is designed for extension. Below are two common scenarios.
Example A: Creating a New Memory Retriever
Scenario: Create a new retriever node called FirstFiveMemories that always gets the first five memories of a
conversation to provide context on its origin.
Implement the Logic (in the Retriever): Open Middleware/services/memory_service.py and add the new method.
# In MemoryService class
def get_first_five_memories(self, discussion_id: str) -> str:
filepath = get_discussion_memory_file_path(discussion_id)
hashed_chunks = read_chunks_with_hashes(filepath)
if not hashed_chunks:
return "No memories have been generated yet"
chunks = extract_text_blocks_from_hashed_chunks(hashed_chunks)
return '--ChunkBreak--'.join(chunks[:5])
Create and Route the Node Type (in the Router): Open Middleware/workflows/handlers/impl/memory_node_handler.py
and add the new type to the handle method's router logic.
# In MemoryNodeHandler.handle()
# ...
elif node_type == "FirstFiveMemories":
return self.memory_service.get_first_five_memories(context.discussion_id)
# ...
Register the Node Type (in the Registrar): Open Middleware/workflows/managers/workflow_manager.py and add the
new node_type to the node_handlers dictionary in the constructor.
# In WorkflowManager.__init__()
self.node_handlers = {
# ...
"FirstFiveMemories": memory_node_handler,
# ...
}
Example B: Using a Workflow for Memory Creation
Scenario: Use a multi-step workflow to generate higher-quality vector-based memories. The first step will identify
key topics, and the second step will write a structured JSON memory object for each topic.
Configure the Memory System: In your discussion ID config file, set useVectorForQualityMemory to true and
specify the name of the workflow to run.
{
"useVectorForQualityMemory": true,
"vectorMemoryWorkflowName": "my-vector-memory-workflow",
"fileMemoryWorkflowName": "my-file-memory-workflow"
}
Understand the Injected Context: The system automatically passes context into your workflow as scoped_inputs,
which are then available in the workflow prompts as {agent1Input}, {agent2Input}, etc. The order is critical.
- For a file-based memory workflow:
{agent1Input}: The raw text chunk to be summarized.
{agent2Input}: The most recent memory chunks.
{agent3Input}: The full history of all memory chunks.
{agent4Input}: The current rolling chat summary.
- For a vector-based memory workflow:
{agent1Input}: The raw text chunk to be processed.
Create the Memory Workflow: Create a new workflow file (e.g.,
Public/Configs/Workflows/my-vector-memory-workflow.json). The final output of this workflow (from the last node
where returnToUser is true) will become the new memory.
Workflow Definition (.../my-vector-memory-workflow.json):
[
{
"id": "agent1",
"type": "Standard",
"prompt": "You are a topic analyzer. Read the following conversation chunk and list the 3 most important topics discussed. Chunk: {agent1Input}",
"returnToUser": false
},
{
"id": "agent2",
"type": "Standard",
"prompt": "You are a summarizer. The original conversation chunk is: {agent1Input}. The topics identified were: {agent1Output}. Generate a JSON array of memory objects, with one object for each topic. Each object must have 'title', 'summary', 'entities', and 'key_phrases' fields.",
"returnToUser": true
}
]
Final Output from Workflow: For a vector memory workflow, this must be a JSON string representing a single
object or an array of objects. For example:
[
{
"title": "Topic A Summary",
"summary": "A detailed summary about the first topic discussed in the chunk.",
"entities": ["Entity1", "Entity2"],
"key_phrases": ["key phrase 1", "key phrase 2"]
},
{
"title": "Topic B Summary",
"summary": "A detailed summary about the second topic discussed in the chunk.",
"entities": ["Entity3"],
"key_phrases": ["key phrase 3"]
}
]
1---2name: 1-core-concepts-and-architecture-33description: This guide provides a deep dive into the architecture and implementation of the conversational memory features within WilmerAI.4---5### **Developer Guide: WilmerAI Memory System**67This guide provides a deep dive into the architecture and implementation of the conversational memory features within8WilmerAI. It has been updated to reflect the powerful workflow-based memory generation system, the specifics of the9vector search implementation, and the complete data schemas as verified by the current codebase.1011-----1213## 1\. Core Concepts & Architecture1415WilmerAI's memory system is a sophisticated, multi-layered feature set implemented through specialized nodes within the16workflow engine. This design allows for different types of memory operations—creation, retrieval, and summarization—to17be strategically placed within workflows.1819#### **Key Architectural Principles:**2021* **Activation via `discussionId`**: All persistent, stateful memory features are tied to a `discussionId`. Its presence22 activates persistent storage mechanisms; its absence causes memory nodes to fall back to stateless, in-memory23 operations using the current chat history.2425* **Separation of Concerns: Creators vs. Retrievers**: The system is fundamentally split into two categories of26 operations, handled by distinct services:2728 * **Memory Creation (Write)**: These are computationally expensive processes that analyze conversation history,29 generate new summarized memories, and write them to persistent storage. This is the exclusive responsibility of30 the **`SlowButQualityRAGTool`**. This tool can generate memories either by executing a full sub-workflow for31 complex logic or by making a direct LLM call.32 * **Memory Retrieval (Read)**: These are inexpensive, fast operations that read from storage or the current chat33 history to provide context for an LLM. This is the primary responsibility of the **`MemoryService`**.3435* **Node-Based Implementation**: Each memory operation is defined as a specific `node_type` in a workflow's JSON36 configuration. This gives developers explicit control over when to perform expensive write operations versus cheap37 read operations.3839* **Centralized Routing and Registration**:4041 * The **`MemoryNodeHandler`** acts as a central router, directing requests for different memory `node_type`s to the42 appropriate service (`MemoryService` or `SlowButQualityRAGTool`).43 * The **`WorkflowManager`** acts as the system registrar where new memory `node_type`s must be mapped to the44 `MemoryNodeHandler`.4546* **Persistent Storage**: When a `discussionId` is active, the system maintains state in a user-specific directory using47 a set of discussion-specific files:4849 1. **Memory File (`<id>_memories.jsonl`)**: Stores discrete, summarized chunks of the conversation for file-based50 memory. Each chunk is saved with a hash of the last message it's based on, creating a traceable, append-only51 ledger.52 2. **Chat Summary File (`<id>_summary.jsonl`)**: Stores a single, continuously updated "rolling summary" of the53 entire conversation. It's linked via a hash to the last memory chunk from the memory file that it incorporated.54 3. **Vector Memory Database (`<id>_vector_memory.db`)**: A **discussion-specific SQLite database** created on-demand55 for each `discussionId`. It uses the FTS5 extension for powerful, weighted, full-text search across two main56 tables:57 * **`memories` table**: Stores the ground-truth data. For vector memories, the `memory_text` column holds the *58 *LLM-generated summary**, not the raw conversation chunk. It also stores the full `metadata_json`.59 * **`memories_fts` table**: A virtual table that indexes the metadata for fast searching. The indexed columns60 are `title`, `summary`, `entities`, `key_phrases`, and the original `memory_text`. Search relevance is61 determined by the **`bm25`** ranking function.62 * **Recency Scoring**: The database connection is initialized with a custom SQL function, `recency_score`, which63 can calculate a time-decay boost for memories. While the default search query uses only `bm25` ranking, this64 function is available for developers to implement time-sensitive search ranking logic.65 4. **Vector Memory Tracker (`vector_memory_tracker` table)**: Located inside the `<id>_vector_memory.db`, this table66 stores the hash of the last message processed for vector memory creation. This crucial feature prevents the67 system from re-processing the same conversation history on subsequent runs.6869-----7071## 2\. Anatomy of Memory Node Types7273Understanding the creator/retriever pattern clarifies the role of each node type.7475#### **Memory Creation & Persistence (Write)**7677These nodes perform the "heavy lifting" of creating and saving memories. They are powered by **`SlowButQualityRAGTool`78**.7980* **`QualityMemory`**: This is the primary **memory creator** node. It's designed to be run periodically in a workflow81 to keep persistent memory up-to-date.82 * **Process Flow**:83 1. The node triggers `SlowButQualityRAGTool.handle_discussion_id_flow`.84 2. The tool checks if new memories are needed by comparing the current conversation history against the last85 processed point.86 3. **(Correction Start)** It checks the `useVectorForQualityMemory` flag from the config to determine its87 behavior: if `true`, it generates **vector memories**; if `false`, it generates **file-based memories**. This88 flag controls the behavior of the `QualityMemory` node specifically. Other nodes (like `RecentMemory`) can89 trigger the creation of file-based memories independently of this flag's value. **(Correction End)**90 4. For the chosen memory type, the tool generates memories using one of two methods, based on its configuration:91 * **A) Workflow-Based Generation (Recommended)** ✨: If the config specifies a `fileMemoryWorkflowName` or92 `vectorMemoryWorkflowName`, the tool executes that sub-workflow.93 * **File-Based Workflow**: Passes the raw text chunk, recent memories, full memories, and the current94 chat summary as `scoped_inputs`. The workflow's final output **must be a single summarized text block95 **.96 * **Vector-Based Workflow**: Passes only the raw text chunk as a `scoped_input`. The workflow's final97 output **must be a JSON string representing a single object or an array of objects**.98 * **B) Direct LLM Call (Legacy)**: If a workflow name is not provided, the system falls back to a direct LLM99 call using prompts defined in the config.100101#### **Memory Retrieval (Read)**102103These nodes perform fast, inexpensive "read" operations and are powered by **`MemoryService`**.104105* **`RecentMemory` & `RecentMemorySummarizerTool`**: The primary **file-based memory retriever**.106107 * **With `discussionId` (Stateful)**: Calls `MemoryService.get_recent_memories`, which reads the last `N` memory108 chunks from `<id>_memories.jsonl`.109 * **Without `discussionId` (Stateless)**: Falls back to an in-memory operation, grabbing the last `N` turns from the110 current `messages` list.111112* **`VectorMemorySearch`**: The primary **RAG search node**.113114 * **Purpose**: Performs a highly relevant, keyword-based search against the discussion-specific vector memory115 database.116 * **Process Flow**: This node takes a string of keywords. The keywords **must be separated by semicolons (`;`)**.117 The `MemoryService` calls `vector_db_utils.search_memories_by_keyword`.118 * **Keyword Handling**: Internally, this process is robust:119 * **Sanitization**: Each keyword is passed through the `_sanitize_fts5_term` function, which wraps it in double120 quotes (`"`) to handle multi-word phrases and prevent FTS5 syntax errors.121 * **Query Construction**: The sanitized terms are combined into a `MATCH` query using **`OR` logic**.122 * **Limits**: The system truncates the keyword list to a maximum of `60` (`MAX_KEYWORDS_FOR_SEARCH`) to avoid123 exceeding SQLite's expression depth limit.124 * **Ranking**: The final results are ranked by relevance using the `bm25` algorithm.125126* **`FullChatSummary`**: Retrieves the holistic, rolling summary of the conversation from `<id>_summary.jsonl`.127128-----129130## 3\. Critical Files for Development131132To modify or extend the memory system, you will primarily work with these five files:1331341. **`Middleware/workflows/tools/slow_but_quality_rag_tool.py`**: The heart of **memory creation**. Modify this file to135 change how memories are generated, including the logic for choosing between vector/file and workflow/LLM-call136 methods.1372. **`Middleware/services/memory_service.py`**: The core of **memory retrieval**. Modify this file to change how138 memories are read from `.jsonl` files or retrieved from the vector database.1393. **`Middleware/utilities/vector_db_utils.py`**: The **database abstraction layer**. This contains all logic for140 interacting with the discussion-specific SQLite databases, including table creation, data insertion, FTS5 search141 query construction, and term sanitization.1424. **`Middleware/workflows/handlers/impl/memory_node_handler.py`**: The **central router**. You must update this file's143 `handle` method to route any new `node_type` you create to the correct service or tool.1445. **`Middleware/workflows/managers/workflow_manager.py`**: The **system registrar**. You must register your new145 `node_type` in the `node_handlers` dictionary within this manager's constructor. This manager is also responsible for146 executing the memory generation sub-workflows.147148-----149150## 4\. Memory Generation Triggering Logic151152This section details the internal logic that determines when new file-based memories are generated. Understanding this153is essential for debugging memory timing issues.154155### Key Configuration Fields156157The triggering logic is governed by four fields in the discussion ID workflow configuration file:158159* **`chunkEstimatedTokenSize`** (default: 1000): The estimated token count threshold. When the accumulated new messages160 reach or exceed this many tokens, memory generation triggers.161* **`maxMessagesBetweenChunks`** (default: 5): The message count threshold. When this many new messages have162 accumulated, memory generation triggers. This threshold only applies when the memory file exists on disk163 (standard mode). It is disabled when the file does not exist (consolidation mode).164* **`maxResponseSizeInTokens`** (default: 400): The maximum number of tokens in the LLM's summarization response.165* **`lookbackStartTurn`** (default: 3): The number of most recent conversation turns to exclude from memory processing.166 This prevents recent, potentially incomplete exchanges (such as frontend command messages) from being included in167 memories.168169### Triggering Modes170171The system operates in two modes depending on whether the memory file exists on disk:172173#### Standard Mode (Memory File Exists)174175Both thresholds are active. Memory generation triggers when **either** threshold is reached first:176177* `estimated_tokens >= chunkEstimatedTokenSize`, OR178* `message_count >= maxMessagesBetweenChunks`179180This "whichever comes first" logic allows the system to generate memories based on either volume (tokens) or frequency181(messages), depending on which condition is met first.182183The memory file does not need to contain any chunks for standard mode to apply — it just needs to exist on disk. In a184new conversation, the file is created (as an empty `[]`) on the first check (around message 3, after the early return185for conversations with fewer than 3 messages). From that point on, standard mode applies and both thresholds are active.186187#### Consolidation Mode (Memory File Does Not Exist)188189Only the token threshold applies. The message count threshold is disabled. This mode activates only when the memory file190has been deleted from disk — for example, when a user wants to regenerate memories with a larger chunk size to produce191fewer, larger chunks.192193### Implementation Details194195The triggering logic lives in `SlowButQualityRAGTool.handle_discussion_id_flow()`, within the file-based memory path196(the `else` branch after the `use_vector_memory_from_config` check).1971981. **File existence check**: Before reading, the system checks `os.path.exists(filepath)` to determine whether the199 memory file exists on disk. This check must happen before `read_chunks_with_hashes()`, which creates the file as an200 empty `[]` if it does not exist (see "File Creation Side-Effect" below).2012. **New message detection**: The method uses hash-based tracking (`find_last_matching_hash_message`) to identify202 messages that have been added since the last memory generation run. Only unprocessed messages are considered. If the203 file has no chunks (new chat or freshly created file), all messages minus the lookback are treated as new.2043. **Token estimation**: The content of all new messages is joined and passed through `rough_estimate_token_length()` to205 get an approximate token count. This function intentionally overestimates by using the higher of a word-based estimate206 (1.35 tokens/word) and a character-based estimate (3.5 chars/token), then applying a `safety_margin` multiplier207 (default 1.10). See `text_utils.py` for details.2084. **Trigger evaluation**: The two boolean conditions are evaluated. `trigger_by_messages` is gated on `file_exists`209 (from step 1). Diagnostic logging reports the exact counts, thresholds, and which trigger(s) fired (or why none did).210211### File Creation Side-Effect212213The `read_chunks_with_hashes()` function calls `ensure_json_file_exists()`, which creates the memory file as an empty214JSON array (`[]`) if it does not exist on disk. This is why the `file_exists` check must happen before that call —215otherwise the file would always appear to exist. The side-effect is intentional: it means that on the first memory216check of a new conversation (around message 3), the file is created, and from message 4 onward the system operates in217standard mode with both thresholds active.218219-----220221## 5\. How to Add a New Memory Feature222223The system is designed for extension. Below are two common scenarios.224225#### **Example A: Creating a New Memory *Retriever***226227**Scenario**: Create a new retriever node called `FirstFiveMemories` that always gets the *first* five memories of a228conversation to provide context on its origin.2292301. **Implement the Logic (in the Retriever)**: Open `Middleware/services/memory_service.py` and add the new method.231232 ```python233 # In MemoryService class234 def get_first_five_memories(self, discussion_id: str) -> str:235 filepath = get_discussion_memory_file_path(discussion_id)236 hashed_chunks = read_chunks_with_hashes(filepath)237 if not hashed_chunks:238 return "No memories have been generated yet"239240 chunks = extract_text_blocks_from_hashed_chunks(hashed_chunks)241 return '--ChunkBreak--'.join(chunks[:5])242 ```2432442. **Create and Route the Node Type (in the Router)**: Open `Middleware/workflows/handlers/impl/memory_node_handler.py`245 and add the new type to the `handle` method's router logic.246247 ```python248 # In MemoryNodeHandler.handle()249 # ...250 elif node_type == "FirstFiveMemories":251 return self.memory_service.get_first_five_memories(context.discussion_id)252 # ...253 ```2542553. **Register the Node Type (in the Registrar)**: Open `Middleware/workflows/managers/workflow_manager.py` and add the256 new `node_type` to the `node_handlers` dictionary in the constructor.257258 ```python259 # In WorkflowManager.__init__()260 self.node_handlers = {261 # ...262 "FirstFiveMemories": memory_node_handler,263 # ...264 }265 ```266267#### **Example B: Using a Workflow for Memory *Creation***268269**Scenario**: Use a multi-step workflow to generate higher-quality vector-based memories. The first step will identify270key topics, and the second step will write a structured JSON memory object for each topic.2712721. **Configure the Memory System**: In your discussion ID config file, set `useVectorForQualityMemory` to `true` and273 specify the name of the workflow to run.274275 ```json276 {277 "useVectorForQualityMemory": true,278 "vectorMemoryWorkflowName": "my-vector-memory-workflow",279 "fileMemoryWorkflowName": "my-file-memory-workflow"280 }281 ```2822832. **Understand the Injected Context**: The system automatically passes context into your workflow as `scoped_inputs`,284 which are then available in the workflow prompts as `{agent1Input}`, `{agent2Input}`, etc. The order is critical.285286 * For a **file-based** memory workflow:287 * `{agent1Input}`: The raw text chunk to be summarized.288 * `{agent2Input}`: The most recent memory chunks.289 * `{agent3Input}`: The full history of all memory chunks.290 * `{agent4Input}`: The current rolling chat summary.291 * For a **vector-based** memory workflow:292 * `{agent1Input}`: The raw text chunk to be processed.2932943. **Create the Memory Workflow**: Create a new workflow file (e.g.,295 `Public/Configs/Workflows/my-vector-memory-workflow.json`). The final output of this workflow (from the last node296 where `returnToUser` is `true`) will become the new memory.297298 *Workflow Definition (`.../my-vector-memory-workflow.json`):*299300 ```json301 [302 {303 "id": "agent1",304 "type": "Standard",305 "prompt": "You are a topic analyzer. Read the following conversation chunk and list the 3 most important topics discussed. Chunk: {agent1Input}",306 "returnToUser": false307 },308 {309 "id": "agent2",310 "type": "Standard",311 "prompt": "You are a summarizer. The original conversation chunk is: {agent1Input}. The topics identified were: {agent1Output}. Generate a JSON array of memory objects, with one object for each topic. Each object must have 'title', 'summary', 'entities', and 'key_phrases' fields.",312 "returnToUser": true313 }314 ]315 ```316317 *Final Output from Workflow:* For a vector memory workflow, this **must be a JSON string** representing a single318 object or an array of objects. For example:319320 ```json321 [322 {323 "title": "Topic A Summary",324 "summary": "A detailed summary about the first topic discussed in the chunk.",325 "entities": ["Entity1", "Entity2"],326 "key_phrases": ["key phrase 1", "key phrase 2"]327 },328 {329 "title": "Topic B Summary",330 "summary": "A detailed summary about the second topic discussed in the chunk.",331 "entities": ["Entity3"],332 "key_phrases": ["key phrase 3"]333 }334 ]335 ```