Memory Skill (Direct Execution)
Direct execution skill for memory vault management. Handles memory creation, similarity search, classification, and index maintenance through content mapping, MCP-based deduplication, and three memory operations (UPDATE, EXTEND, CREATE).
MANDATORY INTERACTIVE REQUIREMENT -- DO NOT SKIP:
- STOP at Step 4 and call AskUserQuestion to show files. Write NOTHING to disk until user responds.
- STOP at Memory Search and call AskUserQuestion for each segment. Write NOTHING to disk until user responds.
- These are not optional. Running autonomously without user input is a critical failure.
Context References
Reference (do not load eagerly):
- Path:
@.memory/30-Templates/memory-template.md- Memory template - Path:
@.memory/20-Indices/index.md- Memory index - Path:
@.memory/memory-index.json- Machine-queryable memory index - Path:
@.opencode/context/project/memory/learn-usage.md- Usage guide
Execution Modes
| Mode | Input | Description |
|---|---|---|
text |
Text content | Add quoted text as memory |
file |
File path | Add single file content as memory |
directory |
Directory path | Scan directory for learnable content |
task |
Task number | Review task artifacts and create memories |
All non-task modes flow through: Content Mapping -> Memory Search -> Memory Operations
Content Mapping
Content mapping is the intermediate representation between input acquisition and memory operations. It segments input into topic-aligned chunks that can be matched against existing memories.
Content Map Data Structure
{
"source": {
"type": "text|file|directory",
"path": "/path/to/input",
"total_tokens": 2500
},
"segments": [
{
"id": "seg-001",
"topic": "python/libs/requests",
"source_file": "/path/to/file.md",
"source_lines": "15-42",
"summary": "HTTP request retry pattern with backoff",
"estimated_tokens": 350,
"key_terms": ["requests", "retry", "backoff", "session", "timeout"]
}
]
}
Field Descriptions
| Field | Type | Description |
|---|---|---|
id |
string | Unique segment identifier (seg-NNN) |
topic |
string | Inferred topic path (slash-separated hierarchy) |
source_file |
string | Original file path (for file/directory modes) |
source_lines |
string | Line range in source file (e.g., "15-42") |
summary |
string | 1-2 sentence summary of segment content |
estimated_tokens |
number | Approximate token count for this segment |
key_terms |
array | 3-5 significant terms for matching |
Segmentation Algorithms
Structured Files (Markdown)
Split at heading boundaries:
1. Identify all headings (# ## ### ####)
2. Each heading starts a new segment
3. Segment includes all content until next same-or-higher level heading
4. Top-level content before first heading becomes its own segment
Structured Files (Code)
Split at blank-line-separated blocks:
1. Identify function/class definitions
2. Group related comments with their definitions
3. Separate standalone comment blocks as documentation segments
4. Keep import/require blocks together
Unstructured Text
Split at paragraph boundaries with topic grouping:
1. Split at double-newline (paragraph boundaries)
2. Group adjacent paragraphs with keyword overlap >40%
3. Single-sentence paragraphs merge with adjacent
Directory Input
Each file becomes an initial segment, then large files are split:
1. Each file is an initial segment
2. Files >800 tokens are split at section boundaries
3. Files <100 tokens are candidates for merging with related files
Small-Input Bypass
Inputs under 500 tokens skip segmentation and become a single segment:
if total_tokens < 500:
segments = [{
"id": "seg-001",
"topic": inferred_topic,
"summary": first_line_or_60_chars,
"estimated_tokens": total_tokens,
"key_terms": extract_keywords(content, 5)
}]
Segment Size Guidelines
| Condition | Action |
|---|---|
| Segment <100 tokens | Merge with adjacent same-topic segment |
| Segment 200-500 tokens | Ideal size, no action |
| Segment >800 tokens | Split at next heading/paragraph boundary |
Key Term Extraction
Extract 3-5 significant terms per segment:
1. Remove stop words (the, a, is, are, etc.)
2. Extract nouns and technical terms (>4 characters)
3. Prioritize: proper nouns > technical terms > common nouns
4. Deduplicate (case-insensitive)
5. Return top 5 by frequency within segment
Memory Search
After content mapping, each segment is matched against existing memories to determine the appropriate operation (UPDATE, EXTEND, or CREATE).
MCP Search Path
When MCP server is available, use the execute pattern:
For each segment in content_map.segments:
query = segment.key_terms.join(" ")
results = execute("search", {
"query": query,
"vault": ".memory",
"limit": 5
})
Grep Fallback Path
When MCP is unavailable, use keyword-based file search:
# For each segment
for keyword in $key_terms; do
grep -l -i "$keyword" .memory/10-Memories/*.md 2>/dev/null
done | sort | uniq -c | sort -rn | head -5
Overlap Scoring
Score keyword overlap between segment and each matching memory:
overlap_score = |segment_terms intersect memory_terms| / |segment_terms|
Where:
- segment_terms = segment.key_terms
- memory_terms = keywords extracted from memory content (same algorithm)
Classification Thresholds
| Overlap Score | Classification | Action |
|---|---|---|
| >60% | HIGH | UPDATE - Replace memory content |
| 30-60% | MEDIUM | EXTEND - Append new section |
| <30% | LOW | CREATE - New memory |
Search Result Presentation -- MANDATORY STOP
YOU MUST call AskUserQuestion for EACH segment before writing anything. Do NOT infer what the user wants. Do NOT skip segments. Do NOT write memory files without explicit user confirmation per segment.
Present each segment with related memories via AskUserQuestion:
Segment: {segment.summary}
Topic: {segment.topic}
Key terms: {segment.key_terms.join(", ")}
Related Memories:
1. MEM-requests-retry-patterns (72% overlap) -> Recommended: UPDATE
2. MEM-python-http-patterns (45% overlap) -> Recommended: EXTEND
3. MEM-api-error-handling (18% overlap) -> Recommended: CREATE (no strong match)
What would you like to do with this segment?
[ ] UPDATE MEM-requests-retry-patterns (replace content)
[ ] EXTEND MEM-python-http-patterns (append section)
[ ] CREATE new memory
[ ] SKIP - don't save this segment
Interactive Override
Users can override any recommendation:
- Change UPDATE to CREATE (preserve existing, create duplicate)
- Change EXTEND to UPDATE (replace instead of append)
- Skip any segment
- Merge segments before processing (combine into single memory)
Memory Operations
Three distinct operations for memory management:
UPDATE Operation
Replace memory content while preserving structure:
1. Read existing memory file
2. Preserve frontmatter: created (original), tags, topic
3. Update frontmatter: modified = today
4. Move current content to ## History section with date marker
5. Replace main content with new segment content
6. Preserve ## Connections section
7. Write updated memory
Template for UPDATE:
---
title: "{new_title_from_segment}"
created: {original_created}
tags: {merged_tags}
topic: "{existing_or_updated_topic}"
source: "{new_source}"
modified: {today}
---
# {new_title}
{new_content_from_segment}
## History
### Previous Version ({original_created})
{previous_content}
## Connections
{preserved_connections}
EXTEND Operation
Append new dated section without modifying existing content:
1. Read existing memory file
2. Find insertion point (before ## Connections, or end of file)
3. Add dated extension section
4. Update frontmatter: modified = today
5. Optionally update tags if new topics introduced
6. Write updated memory
Template for EXTEND:
## Extension ({today})
**Source**: {segment.source_file}
{segment_content}
CREATE Operation
Generate new memory from segment:
1. Generate semantic slug from topic and title:
generate_slug() {
local topic="$1"
local title="$2"
local base=""
# Priority 1: Topic path (most specific segment)
if [ -n "$topic" ]; then
base=$(echo "$topic" | rev | cut -d'/' -f1 | rev)
fi
# Priority 2: First 2-3 words of title
local title_slug=$(echo "$title" | tr '[:upper:]' '[:lower:]' | \
sed 's/[^a-z0-9 ]/-/g' | tr ' ' '-' | \
cut -d'-' -f1-3 | sed 's/-$//')
# Combine
if [ -n "$base" ]; then
slug="${base}-${title_slug}"
else
slug="$title_slug"
fi
# Sanitize and truncate to 50 chars
slug=$(echo "$slug" | sed 's/--*/-/g' | sed 's/^-//' | sed 's/-$//' | cut -c1-50)
# Handle collision - NOTE: MEM- prefix preserved for grep discoverability
local final_slug="$slug"
local counter=2
while [ -f ".memory/10-Memories/MEM-${final_slug}.md" ]; do
final_slug="${slug}-${counter}"
counter=$((counter + 1))
done
echo "$final_slug"
}
slug=$(generate_slug "$topic" "$title")
filename="MEM-${slug}.md"
2. Apply memory template with all fields
3. Infer and apply topic
4. Add to index (both category and topic sections)
5. Write new memory file
Template for CREATE:
---
title: "{segment.summary}"
created: {today}
tags: {inferred_tags}
topic: "{segment.topic}"
source: "{segment.source_file or 'user input'}"
modified: {today}
keywords: {segment.key_terms}
summary: "{segment.summary}"
retrieval_count: 0
last_retrieved:
---
# {segment.summary}
{segment_content}
## Connections
<!-- Add links to related memories using [[filename]] syntax -->
Note: The MEM- prefix is preserved for grep discoverability (grep -r "MEM-" .memory/). Filenames follow the pattern MEM-{semantic-slug}.md (e.g., MEM-requests-retry-patterns.md).
Topic Inference
Infer topic using four-source priority:
1. Source directory path (highest priority)
- /project/src/utils/ -> "project/utils"
- /home/user/notes/python/ -> "python"
2. Keyword analysis
- Extract domain indicators: python, requests, http, api
- Map to topic: "python/libs" or "python/patterns"
3. Related memory topics
- If UPDATE/EXTEND: inherit topic from target memory
- If CREATE with high-overlap match: suggest that topic
4. User confirmation/override
- Always present inferred topic for confirmation
- User can modify or create new topic path
Index Maintenance
Note: After each operation, update all three indexes:
index.md,.memory/10-Memories/README.md, andmemory-index.json. See "JSON Index Maintenance" and "Index Regeneration Pattern" below.
After each operation, update both index.md and .memory/10-Memories/README.md:
index.md:
1. Add/update entry in "## By Category" under appropriate tag
2. Add/update entry in "## By Topic" under topic path
3. Update "## Recent Memories" (prepend, keep last 10)
4. Update "## Statistics" counts
.memory/10-Memories/README.md -- regenerate the full file listing:
1. List all MEM-*.md files in the directory (ls .memory/10-Memories/MEM-*.md)
2. For each file, extract: title, topic, tags, created from frontmatter
3. Rewrite README.md with updated count and one entry per memory:
### [MEM-{slug}](MEM-{slug}.md)
**Title**: {title}
**Topic**: {topic}
**Tags**: {tags}
**Created**: {created}
4. Keep "## Navigation" section at the bottom
Index Regeneration Pattern
To avoid concurrent write conflicts, regenerate index.md from filesystem state rather than append:
# 1. List all memory files
memories=$(ls .memory/10-Memories/MEM-*.md 2>/dev/null)
# 2. Extract metadata from each file
for mem in $memories; do
title=$(grep -m1 "^title:" "$mem" | cut -d'"' -f2)
topic=$(grep -m1 "^topic:" "$mem" | cut -d'"' -f2)
created=$(grep -m1 "^created:" "$mem" | cut -d: -f2 | tr -d ' ')
# Store for index generation
done
# 3. Regenerate index.md from extracted data
# Sort by date descending, write complete file
Benefits:
- No append conflicts (complete overwrite)
- Self-healing (missing entries recovered)
- Idempotent (multiple regenerations produce same result)
JSON Index Maintenance
After each CREATE, UPDATE, or EXTEND operation, regenerate .memory/memory-index.json from filesystem state:
# 1. Scan all memory files
memories=$(ls .memory/10-Memories/MEM-*.md 2>/dev/null)
# 2. For each file, extract frontmatter fields
for mem in $memories; do
title=$(grep -m1 "^title:" "$mem" | sed 's/^title: *//' | tr -d '"')
topic=$(grep -m1 "^topic:" "$mem" | sed 's/^topic: *//' | tr -d '"')
created=$(grep -m1 "^created:" "$mem" | sed 's/^created: *//')
modified=$(grep -m1 "^modified:" "$mem" | sed 's/^modified: *//')
keywords=$(grep -m1 "^keywords:" "$mem" | sed 's/^keywords: *//')
summary=$(grep -m1 "^summary:" "$mem" | sed 's/^summary: *//' | tr -d '"')
retrieval_count=$(grep -m1 "^retrieval_count:" "$mem" | sed 's/^retrieval_count: *//')
last_retrieved=$(grep -m1 "^last_retrieved:" "$mem" | sed 's/^last_retrieved: *//')
status=$(grep -m1 "^status:" "$mem" | sed 's/^status: *//')
# Default status to "active" when absent
if [ -z "$status" ]; then status="active"; fi
# Compute token_count: word_count * 1.3
word_count=$(wc -w < "$mem")
token_count=$(echo "$word_count * 1.3" | bc | cut -d. -f1)
# Derive id from filename: MEM-{slug}.md -> MEM-{slug}
id=$(basename "$mem" .md)
# Derive category from first tag
category=$(grep -m1 "^tags:" "$mem" | sed 's/^tags: *\[//' | cut -d, -f1 | tr -d '] ')
done
# 3. Build JSON structure
{
"version": "1.0.0",
"generated_at": "$(date +%Y-%m-%d)",
"entry_count": N,
"total_tokens": sum_of_token_counts,
"entries": [...]
}
# 4. Write to .memory/memory-index.json (complete overwrite)
Schema Fields per Entry:
| Field | Type | Source |
|---|---|---|
id |
string | Filename without .md extension |
path |
string | Relative path from project root |
title |
string | Frontmatter title |
summary |
string | Frontmatter summary |
topic |
string | Frontmatter topic |
category |
string | First tag from frontmatter tags |
keywords |
array | Frontmatter keywords |
token_count |
number | Word count * 1.3, rounded down |
created |
string | Frontmatter created (ISO date) |
modified |
string | Frontmatter modified (ISO date) |
last_retrieved |
string/null | Frontmatter last_retrieved |
retrieval_count |
number | Frontmatter retrieval_count |
status |
string | Frontmatter status (default: "active" when absent; "tombstoned" for purged memories) |
Validate-on-Read
Before using memory-index.json for retrieval or scoring, validate that the index matches the filesystem:
1. List all MEM-*.md files in .memory/10-Memories/
2. List all entry ids in memory-index.json
3. Compare:
- Files on disk not in index -> INDEX STALE (missing entries)
- Index entries with no file on disk -> INDEX STALE (orphaned entries)
- All match -> INDEX VALID
4. If INDEX STALE: regenerate memory-index.json using JSON Index Maintenance procedure
5. If INDEX VALID: proceed with retrieval
This ensures the index is always consistent, even if manual file edits bypass the skill pipeline.
Status Field Handling: During regeneration, the status field is read from each memory's frontmatter. If absent, it defaults to "active". Tombstoned memories (with status: tombstoned in frontmatter) retain their "tombstoned" status in the regenerated index. The tombstoned_at and tombstone_reason fields are also preserved when present.
Task Mode Execution
Task mode has special handling for reviewing existing task artifacts.
Step 1: Locate Task Directory
task_num=$task_number
padded_num=$(printf "%03d" $task_num)
task_dir=$(ls -d specs/${padded_num}_* 2>/dev/null | head -1)
if [ -z "$task_dir" ]; then
task_dir=$(ls -d specs/${task_num}_* 2>/dev/null | head -1)
fi
if [ -z "$task_dir" ]; then
echo "Task directory not found: specs/${padded_num}_*"
exit 1
fi
Step 2: Scan Artifacts
artifacts=$(find "$task_dir" -type f -name "*.md" | sort)
if [ -z "$artifacts" ]; then
echo "No artifacts found for task ${task_number}"
exit 1
fi
Step 3: Present Artifact List
Display via AskUserQuestion:
{
"question": "Select artifacts to review for memory extraction:",
"header": "Task Artifacts",
"multiSelect": true,
"options": [
{
"label": "{artifact_1_name}",
"description": "{artifact_1_path}"
}
]
}
Step 4: Process Through Content Mapping
For each selected artifact:
- Read content
- If >500 tokens: run through content mapping (segmentation)
- If <=500 tokens: treat as single segment
- Proceed to Memory Search (Phase 4)
- Proceed to Memory Operations (Phase 5)
Step 5: Classification Taxonomy
For task artifacts, also present classification options:
{
"question": "Classify this segment:",
"header": "Classification: {segment.summary}",
"multiSelect": false,
"options": [
{"label": "[TECHNIQUE]", "description": "Reusable method or approach"},
{"label": "[PATTERN]", "description": "Design or implementation pattern"},
{"label": "[CONFIG]", "description": "Configuration or setup knowledge"},
{"label": "[WORKFLOW]", "description": "Process or procedure"},
{"label": "[INSIGHT]", "description": "Key learning or understanding"},
{"label": "[SKIP]", "description": "Not valuable for memory"}
]
}
Step 6: Return Result
{
"status": "completed",
"mode": "task",
"artifacts_reviewed": [...],
"content_map": { ... },
"operations": [
{"type": "CREATE", "memory_id": "MEM-...", "category": "[PATTERN]"}
],
"memories_affected": 3
}
Directory Mode Execution
Directory mode scans a directory tree for learnable content.
Step 1: Recursive Scanning
# Exclusion patterns
EXCLUDES="-path '*/.git' -prune -o -path '*/node_modules' -prune -o -path '*/__pycache__' -prune -o -path '*/.obsidian' -prune"
# Find all files
files=$(find "$directory_path" $EXCLUDES -type f -print | head -250)
Step 2: Two-Tier Text Detection
Tier 1: Extension Whitelist
Recognized text extensions (alphabetized by category):
| Category | Extensions |
|---|---|
| Code | .c, .cpp, .cs, .go, .h, .hpp, .java, .js, .jsx, .kt, .lua, .php, .pl, .py, .r, .rb, .rs, .scala, .sh, .swift, .ts, .tsx, .vim |
| Config | .cfg, .conf, .ini, .json, .toml, .xml, .yaml, .yml |
| Data | .csv, .sql |
| Documentation | .adoc, .asciidoc, .md, .org, .rdoc, .rst, .tex, .txt |
| Web | .css, .htm, .html, .less, .sass, .scss, .svg |
| Scripting | .fnl, .janet, .nix |
Tier 2: MIME-Type Fallback
For files without recognized extensions:
mime=$(file --mime-type -b "$file")
if [[ "$mime" == text/* ]]; then
# Include file
fi
Step 3: Size Limits
# Per-file limit
if [ $(stat -c%s "$file") -gt 102400 ]; then
echo "Skipping large file: $file (>100KB)"
continue
fi
# Warning at 50 files
if [ ${#files[@]} -gt 50 ]; then
echo "Warning: ${#files[@]} files found. Consider narrowing scope."
fi
# Hard limit at 200 files
if [ ${#files[@]} -gt 200 ]; then
echo "Error: Too many files (${#files[@]}). Maximum is 200."
echo "Narrow your path or use file mode for specific files."
exit 1
fi
Step 4: File Selection (Paginated) -- MANDATORY STOP
YOU MUST call AskUserQuestion here. Do NOT skip to Step 5. Do NOT process any files until the user has made their selection.
Present files in pages of 10 to avoid overwhelming the display. Accumulate selections across all pages before processing.
selected_files = []
page_size = 10
total_files = len(files)
page = 0
while page * page_size < total_files:
start = page * page_size
end = min(start + page_size, total_files)
page_files = files[start:end]
remaining = total_files - end
page_num = page + 1
total_pages = ceil(total_files / page_size)
# Build options for this page
options = [{"label": relative_path, "description": file_size} for each file in page_files]
# Add navigation options at the bottom
if remaining > 0:
options.append({"label": "--- Continue to next page ---", "description": f"{remaining} more files remaining"})
AskUserQuestion({
"question": f"Select files to include (page {page_num}/{total_pages}, showing {start+1}-{end} of {total_files}):",
"header": f"Directory Scan: {directory_path}",
"multiSelect": true,
"options": options
})
# Add any selected files (excluding the navigation option) to accumulated list
selected_files.extend(user_selections excluding navigation option)
# If user selected "Continue to next page" OR there are more pages, advance
# If user did NOT select "Continue to next page" on the last page, stop
if "--- Continue to next page ---" not in user_selections and remaining > 0:
# User is done selecting (didn't ask for more)
break
page += 1
# After all pages processed, confirm total selection
if len(selected_files) == 0:
print("No files selected. Exiting.")
exit
Example page 1 of 3:
{
"question": "Select files to include (page 1/3, showing 1-10 of 28):",
"header": "Directory Scan: /home/user/project/",
"multiSelect": true,
"options": [
{"label": "README.md", "description": "4.1KB"},
{"label": "src/main.lua", "description": "2.3KB"},
{"label": "--- Continue to next page ---", "description": "18 more files remaining"}
]
}
Step 5: Route Through Pipeline
For each selected file:
- Read file content
- Run through content mapping (directory-type segmentation)
- Route segments through memory search
- Route through memory operations
- Update index
Step 6: Return Result
{
"status": "completed",
"mode": "directory",
"files_scanned": 45,
"files_selected": 12,
"content_map": { ... },
"operations": [...],
"memories_affected": 8
}
Text Mode Execution
Step 1: Parse Input
content="$text_content"
source="user input"
Step 2: Content Mapping
For text >500 tokens, segment at paragraph boundaries:
1. Split at double-newline
2. Group related paragraphs
3. Generate single content map
For text <500 tokens, create single segment.
Step 3: Memory Search & Operations
Route through standard memory search and operations pipeline.
Step 4: Return Result
{
"status": "completed",
"mode": "text",
"content_map": { ... },
"operations": [...],
"memories_affected": 1
}
File Mode Execution
Step 1: Read File
if [ ! -f "$file_path" ]; then
echo "File not found: $file_path"
exit 1
fi
content=$(cat "$file_path")
source="file: $file_path"
Step 2: Content Mapping
Apply structured or unstructured segmentation based on file type.
Step 3: Memory Search & Operations
Route through standard pipeline.
Step 4: Return Result
{
"status": "completed",
"mode": "file",
"file_path": "...",
"content_map": { ... },
"operations": [...],
"memories_affected": 2
}
Error Handling
No Content Provided
Usage: /learn <text or file path or directory> OR /learn --task N
File Not Found
File not found: {path}
Directory Not Found
Directory not found: {path}
Empty Directory
No text files found in: {path}
Too Many Files
Too many files ({N}). Maximum is 200.
Narrow your path or use file mode for specific files.
Task Directory Not Found
Task directory not found: specs/{NNN}_*
User Cancels
Memory operation cancelled. No files created.
All Content Skipped
No memories created (all content skipped)
MCP Unavailable
MCP search unavailable. Using grep-based fallback.
Git Commit (Postflight)
After successful memory operations:
git add .memory/
git commit -m "memory: add/update ${memories_affected} memories
Session: ${session_id}
Mode: distill
Memory vault distillation: scoring, health reporting, and maintenance operations. Invoked by /distill command with mode=distill.
Prerequisites
Validate-on-Read: Before scoring, run the validate-on-read procedure from the "Validate-on-Read" section above to ensure memory-index.json is consistent with the filesystem. If stale, regenerate using the "JSON Index Maintenance" procedure before proceeding.
Sub-Mode Dispatch
| Sub-Mode | Description | Status |
|---|---|---|
report |
Generate health report with scoring | Available |
purge |
Tombstone stale/zero-retrieval memories | Available |
merge |
Combine memories with duplicate score > 0.6 | Available |
compress |
Summarize memories with size penalty > 0.5 | Available |
refine |
Improve memory quality (keywords, tags) | Available |
gc |
Hard-delete tombstoned memories past grace period | Available |
auto |
Automated distillation (Tier 1 refine only) | Available |
All sub-modes are now available. No placeholder responses needed.
Scoring Engine
The scoring engine computes a composite maintenance score for each memory in the vault. Higher scores indicate memories that are better candidates for maintenance operations.
Input
Read all entries from .memory/memory-index.json after validate-on-read. Each entry provides:
created(ISO date)modified(ISO date)last_retrieved(ISO date or null)retrieval_count(number)token_count(number)keywords(array of strings)
Component 1: Staleness Score (weight: 0.3)
Measures how long since the memory was last useful.
days_since_last = days_between(today, last_retrieved or created)
staleness = min(1.0, days_since_last / 90)
# FSRS adjustment: reduce staleness for actively retrieved old memories
if retrieval_count > 0 AND days_since_created > 60:
staleness = max(0, staleness - 0.3)
- Range: 0.0 (fresh) to 1.0 (90+ days stale)
- FSRS adjustment rewards memories that have proven useful over time
Component 2: Zero-Retrieval Penalty (weight: 0.25)
Penalizes memories that have never been retrieved after a grace period.
if retrieval_count == 0 AND days_since_created > 30:
zero_retrieval = 1.0
else:
zero_retrieval = 0.0
- Binary: 0.0 (has retrievals or too new) or 1.0 (never retrieved, older than 30 days)
Component 3: Size Penalty (weight: 0.2)
Penalizes oversized memories that may benefit from compression.
size_penalty = max(0, (token_count - 600) / 600)
- Range: 0.0 (600 tokens or fewer) to unbounded (linear above 600)
- A 1200-token memory scores 1.0; a 300-token memory scores 0.0
Component 4: Duplicate Score (weight: 0.25)
Measures keyword overlap with the most similar other memory in the vault.
for each other_memory in vault:
overlap = |memory.keywords intersect other_memory.keywords| / |memory.keywords|
duplicate = max(overlap across all other memories)
- Range: 0.0 (no keyword overlap) to 1.0 (complete keyword subset)
- Uses Jaccard-like ratio: intersection size divided by the memory's own keyword count
Composite Score
composite = (staleness * 0.3) + (zero_retrieval * 0.25) + (size_penalty * 0.2) + (duplicate * 0.25)
composite = clamp(composite, 0, 1)
- Weights sum to 1.0 (0.3 + 0.25 + 0.2 + 0.25)
- Range: 0.0 (healthy memory) to 1.0 (strong maintenance candidate)
Topic-Cluster Grouping
Group memories by topic cluster for the health report. The cluster key is the first path segment of the memory's topic field:
cluster_key = topic.split("/")[0]
# Example:
# topic "python/libs/requests" -> cluster "python"
# topic "lua/patterns" -> cluster "lua"
# topic "" or null -> cluster "uncategorized"
Maintenance Candidate Classification
Based on composite scores, classify each memory:
| Composite Score | Classification | Recommended Action |
|---|---|---|
| >= 0.7 | Purge candidate | Remove (--purge) |
| >= 0.5 | Merge/compress candidate | Merge duplicates (--merge) or compress (--compress) |
| >= 0.3 | Review candidate | May benefit from refinement (--refine) |
| < 0.3 | Healthy | No action needed |
Additionally, flag specific conditions:
duplicate > 0.6-> Merge candidate regardless of compositesize_penalty > 0.5-> Compress candidate regardless of compositezero_retrieval == 1.0-> Review for relevance
Health Report Template
The report sub-mode generates a formatted health report displayed to the user. Template:
## Memory Vault Health Report
**Generated**: {today}
**Vault**: .memory/
---
### Overview
| Metric | Value |
|--------|-------|
| Total memories | {total_count} |
| Total tokens | {total_tokens} |
| Average tokens/memory | {avg_tokens} |
| Oldest memory | {oldest_date} ({oldest_id}) |
| Newest memory | {newest_date} ({newest_id}) |
---
### Category Distribution
| Category | Count | Tokens | Avg Score |
|----------|-------|--------|-----------|
| {category_1} | {count} | {tokens} | {avg_composite} |
| {category_2} | {count} | {tokens} | {avg_composite} |
| ... | ... | ... | ... |
---
### Topic Clusters
| Cluster | Memories | Avg Staleness | Avg Duplicate |
|---------|----------|---------------|---------------|
| {cluster_1} | {count} | {avg_staleness} | {avg_duplicate} |
| {cluster_2} | {count} | {avg_staleness} | {avg_duplicate} |
| ... | ... | ... | ... |
---
### Retrieval Statistics
| Metric | Value |
|--------|-------|
| Never retrieved | {never_retrieved_count} ({never_retrieved_pct}%) |
| Retrieved 1-3 times | {low_retrieval_count} |
| Retrieved 4+ times | {high_retrieval_count} |
| Most retrieved | {most_retrieved_id} ({most_retrieved_count} times) |
---
### Maintenance Candidates
#### Purge Candidates (score >= 0.7)
{purge_list or "None"}
#### Merge Candidates (duplicate > 0.6)
{merge_list or "None"}
#### Compress Candidates (size > 0.5)
{compress_list or "None"}
#### Review Candidates (score 0.3-0.7)
{review_list or "None"}
---
### Health Score
**Score**: {health_score}/100
**Status**: {status_emoji} {status_label}
Formula: `100 - (purge_count * 3) - (merge_count * 5) - (compress_count * 2)`
| Threshold | Status |
|-----------|--------|
| 80-100 | Healthy |
| 60-79 | Manageable |
| 40-59 | Concerning |
| 0-39 | Critical |
---
### Recommended Actions
{action_list based on candidates found}
Health Score Formula
health_score = 100 - (purge_count * 3) - (merge_count * 5) - (compress_count * 2)
health_score = clamp(health_score, 0, 100)
Where:
purge_count= number of memories with composite score >= 0.7merge_count= number of memories with duplicate score > 0.6compress_count= number of memories with size_penalty > 0.5
Health Status Thresholds
| Score Range | Status | Description |
|---|---|---|
| 80-100 | healthy | Vault is well-maintained |
| 60-79 | manageable | Some maintenance recommended |
| 40-59 | concerning | Significant maintenance needed |
| 0-39 | critical | Urgent maintenance required |
These thresholds mirror repository_health.status vocabulary in state.json.
Sub-Mode: merge
Combine duplicate memories with high keyword overlap. The merge operation identifies pairwise duplicate candidates within topic clusters, presents them for interactive selection, merges content with a keyword superset guarantee, tombstones the absorbed secondary, updates cross-references, and regenerates indexes.
Edge Case Checks
Before candidate identification, validate:
1. Run validate-on-read to ensure memory-index.json is consistent
2. Count non-tombstoned memories (status != "tombstoned" or status absent)
3. If fewer than 2 non-tombstoned memories:
Display: "Merge requires at least 2 active memories. Vault has {count}."
Return early.
Pairwise Keyword Overlap Algorithm
Compute pairwise overlap within each topic cluster:
1. Group non-tombstoned memories by topic cluster:
cluster_key = topic.split("/")[0]
If topic is empty or null: cluster_key = "uncategorized"
2. For each cluster with 2+ memories, compute pairwise overlap:
for each pair (A, B) in cluster:
overlap_ab = |A.keywords intersect B.keywords| / |A.keywords|
overlap_ba = |A.keywords intersect B.keywords| / |B.keywords|
pair_overlap = max(overlap_ab, overlap_ba)
Note: Use max of both asymmetric directions so that a small memory
with all keywords contained in a larger memory is detected.
3. Handle empty keyword arrays:
If either A.keywords or B.keywords is empty: pair_overlap = 0.0
(Cannot merge memories with no keyword basis for comparison)
4. Filter pairs where pair_overlap >= 0.6 (60% threshold)
5. Sort candidate pairs by pair_overlap descending within each cluster
Dry-Run Mode
When --dry-run is set, compute and display candidates without writing any files:
Display per cluster:
## Merge Candidates (Dry Run)
### Cluster: {cluster_key}
| Primary | Secondary | Overlap | Shared Keywords |
|---------|-----------|---------|-----------------|
| {A.id} | {B.id} | {pair_overlap}% | {shared_keywords} |
{total_pairs} merge candidate pair(s) found across {cluster_count} cluster(s).
Run /distill --merge without --dry-run to execute.
Return early after display. No files are modified.
Interactive Selection (AskUserQuestion)
Present merge candidates per topic cluster for user selection:
For each cluster with candidates:
AskUserQuestion({
"question": "Select pairs to merge in cluster '{cluster_key}':",
"header": "Merge Candidates: {cluster_key}",
"multiSelect": true,
"options": [
{
"label": "{A.title} + {B.title}",
"description": "{pair_overlap}% overlap | Shared: {shared_keywords} | Retrievals: {A.retrieval_count}, {B.retrieval_count}"
}
]
})
If no pairs above threshold in any cluster:
Display: "No merge candidates found (no pairs with >= 60% keyword overlap)."
Return early.
If user selects no pairs across all clusters:
Display: "No pairs selected. No merges performed."
Return early.
Primary Determination
For each selected pair, determine which memory is primary (target) and which is secondary (absorbed):
Primary selection rules (first match wins):
1. Higher retrieval_count -> primary
2. If retrieval_count equal: older created date -> primary
3. If both equal: alphabetically first id -> primary (deterministic tiebreaker)
Merged Content Template
The primary memory file is rewritten with merged content:
Frontmatter merging rules:
title: primary.title (unchanged)
created: min(primary.created, secondary.created) -- earliest
modified: today (ISO date)
tags: union(primary.tags, secondary.tags) -- deduplicated
topic: primary.topic (unchanged)
source: primary.source (unchanged)
keywords: union(primary.keywords, secondary.keywords) -- deduplicated, sorted
summary: primary.summary (unchanged)
retrieval_count: primary.retrieval_count + secondary.retrieval_count
last_retrieved: max(primary.last_retrieved, secondary.last_retrieved) -- most recent, skip nulls
token_count: recomputed after merge (word_count * 1.3, rounded down)
status: omit (active is default when absent)
Content structure:
---
{merged frontmatter}
---
# {primary.title}
{primary existing content - everything between title heading and first ## section}
## Merged From {secondary.id}
**Original Title**: {secondary.title}
**Merged**: {today}
**Overlap Score**: {pair_overlap}%
{secondary content - everything between title heading and ## Connections in secondary}
## Connections
{union of both connection sections, with [[{secondary.id}]] references replaced by [[{primary.id}]]}
Keyword Superset Guarantee
CRITICAL INVARIANT: Before writing the merged file, verify:
required_keywords = union(primary.keywords, secondary.keywords)
merged_keywords = merged_frontmatter.keywords
assertion: set(merged_keywords) >= set(required_keywords)
If assertion fails:
Log error: "KEYWORD SUPERSET VIOLATION: missing keywords: {required - merged}"
Abort this merge pair (do not write file)
Preserve both original files unchanged
Continue with remaining pairs
Report violation in operation summary
This guarantee ensures no keyword coverage is lost during merging.
Tombstone Application
After successful merge, tombstone the secondary memory:
Add to secondary's frontmatter (preserve all existing fields):
status: tombstoned
tombstoned_at: {today ISO8601}
tombstone_reason: "merged_into:{primary.id}"
Do NOT delete the file.
Do NOT remove from index (index regeneration will include tombstone status).
The tombstone fields are identical to those used by the purge sub-mode:
status: tombstonedtombstoned_at: {ISO8601 date}tombstone_reason: "{reason}"-- for merge, reason is"merged_into:{primary_id}"
Cross-Reference Update
After tombstoning, update wiki-link references across all non-tombstoned memories:
1. Scan all .memory/10-Memories/*.md files
2. For each file that is NOT tombstoned:
Search for [[{secondary.id}]] references
Replace with [[{primary.id}]]
3. Log all replacements: "{file}: replaced [[{secondary.id}]] -> [[{primary.id}]]"
Index Regeneration
After ALL merges in the batch are complete (not after each individual merge):
1. Regenerate memory-index.json using "JSON Index Maintenance" procedure
- Include tombstoned memories with status: "tombstoned"
2. Regenerate index.md using "Index Regeneration Pattern"
- Exclude tombstoned memories from active listings
3. Regenerate .memory/10-Memories/README.md
- Exclude tombstoned memories from the listing
Distill Log Entry
Log each merge operation to .memory/distill-log.json:
{
"id": "distill_{timestamp}",
"timestamp": "ISO8601",
"type": "merge",
"session_id": "sess_...",
"pre_metrics": {
"total_memories": N,
"total_tokens": N,
"health_score": N,
"purge_candidates": N,
"merge_candidates": N,
"compress_candidates": N
},
"post_metrics": {
"total_memories": N,
"total_tokens": N,
"health_score": N,
"purge_candidates": N,
"merge_candidates": N,
"compress_candidates": N
},
"affected_memories": [
{
"primary": "{primary.id}",
"secondary": "{secondary.id}",
"overlap_score": 0.75,
"keywords_before": [5, 4],
"keywords_after": 7,
"keyword_superset_verified": true,
"action": "merged"
}
],
"notes": "Merged {N} pair(s) across {M} cluster(s)"
}
The keywords_before array contains [primary_keyword_count, secondary_keyword_count]. The keywords_after value is the merged keyword count. The keyword_superset_verified boolean confirms the superset guarantee held for this pair.
Sub-Mode: compress
Reduce oversized memories to key points while preserving essential information. The compress operation identifies memories with high size penalty, presents them interactively, generates compressed versions, preserves originals in a History section, and ensures keyword preservation.
Edge Case Checks
Before candidate identification, validate:
1. Run validate-on-read to ensure memory-index.json is consiste
…(truncated)