Deep Research
Core Purpose
Deliver citation-tracked research reports through a structured pipeline with evidence persistence, source identity management, claim-level verification, and progressive context management.
Autonomy Principle: Operate independently. Infer assumptions from context. Only stop for critical errors or incomprehensible queries. Surface high-materiality assumptions explicitly in the Introduction and Methodology rather than silently defaulting.
Decision Tree
Request Analysis
+-- Simple lookup? --> STOP: Use WebSearch
+-- Debugging? --> STOP: Use standard tools
+-- Complex analysis needed? --> CONTINUE
Mode Selection
+-- Initial exploration --> quick (3 phases, 2-5 min)
+-- Standard research --> standard (6 phases, 5-10 min) [DEFAULT]
+-- Critical decision --> deep (8 phases, 10-20 min)
+-- Comprehensive review --> ultradeep (8+ phases, 20-45 min)
Default assumptions: Technical query = technical audience. Comparison = balanced perspective. Trend = recent 1-2 years.
Workflow Overview
| Phase |
Name |
Quick |
Std |
Deep |
Ultra |
| 1 |
SCOPE |
Y |
Y |
Y |
Y |
| 2 |
PLAN |
- |
Y |
Y |
Y |
| 3 |
RETRIEVE |
Y |
Y |
Y |
Y |
| 4 |
TRIANGULATE |
- |
Y |
Y |
Y |
| 4.5 |
OUTLINE REFINEMENT |
- |
Y |
Y |
Y |
| 5 |
SYNTHESIZE |
- |
Y |
Y |
Y |
| 6 |
CRITIQUE |
- |
- |
Y |
Y |
| 7 |
REFINE |
- |
- |
Y |
Y |
| 8 |
PACKAGE |
Y |
Y |
Y |
Y |
Note: Phases 3-5 operate as an evidence loop per section (retrieve → evidence store → refine outline → draft → verify claims → delta-retrieve if needed), not as strict sequential gates.
Execution
On invocation, load relevant reference files:
- Phase 1-7: Load methodology.md for detailed phase instructions
- Phase 8 (Report): Load report-assembly.md for progressive generation
- HTML/PDF output: Load html-generation.md
- Quality checks: Load quality-gates.md
- Long reports (>18K words): Load continuation.md
Templates:
Scripts:
python scripts/validate_report.py --report [path]
python scripts/verify_citations.py --report [path]
python scripts/md_to_html.py [markdown_path]
Output Contract
Required sections:
- Executive Summary (200-400 words)
- Introduction (scope, methodology, assumptions)
- Main Analysis (4-8 findings, 600-2,000 words each, cited)
- Synthesis & Insights (patterns, implications)
- Limitations & Caveats
- Recommendations
- Bibliography (COMPLETE - every citation, no placeholders)
- Methodology Appendix
Output files (all to ~/Documents/[Topic]_Research_[YYYYMMDD]/):
- Markdown (primary source of truth)
sources.jsonl — stable source registry with canonical IDs
evidence.jsonl — append-only evidence store with quotes and locators
claims.jsonl — atomic claim ledger with support status
run_manifest.json — query, mode, assumptions, provider config
- HTML (McKinsey style, auto-opened)
- PDF (professional print, auto-opened)
Quality standards:
- 10+ sources, 3+ per major claim (cluster-independent, not just count)
- All factual claims cited immediately [N] with evidence backing in
evidence.jsonl
- Claim-support verification mandatory: no unsupported factual claims pass delivery
- No placeholders, no fabricated citations
- Prose-first (>=80%), bullets sparingly
When to Use / NOT Use
Use: Comprehensive analysis, technology comparisons, state-of-the-art reviews, multi-perspective investigation, market analysis.
Do NOT use: Simple lookups, debugging, 1-2 search answers, quick time-sensitive queries.
1---2name: deep-research3description: Use when the user needs multi-source research with citation tracking, evidence persistence, and structured report generation. Triggers on "deep research", "comprehensive analysis", "research report", "compare X vs Y", "analyze trends", or "state of the art". Not for simple lookups, debugging, or questions answerable with 1-2 searches.4---5
6# Deep Research
7
8## Core Purpose
9
10Deliver citation-tracked research reports through a structured pipeline with evidence persistence, source identity management, claim-level verification, and progressive context management.
11
12**Autonomy Principle:** Operate independently. Infer assumptions from context. Only stop for critical errors or incomprehensible queries. Surface high-materiality assumptions explicitly in the Introduction and Methodology rather than silently defaulting.
13
14---
15
16## Decision Tree
17
18```
19Request Analysis
20+-- Simple lookup? --> STOP: Use WebSearch
21+-- Debugging? --> STOP: Use standard tools
22+-- Complex analysis needed? --> CONTINUE
23
24Mode Selection
25+-- Initial exploration --> quick (3 phases, 2-5 min)
26+-- Standard research --> standard (6 phases, 5-10 min) [DEFAULT]
27+-- Critical decision --> deep (8 phases, 10-20 min)
28+-- Comprehensive review --> ultradeep (8+ phases, 20-45 min)
29```
30
31**Default assumptions:** Technical query = technical audience. Comparison = balanced perspective. Trend = recent 1-2 years.
32
33---
34
35## Workflow Overview
36
37| Phase | Name | Quick | Std | Deep | Ultra |
38|-------|------|-------|-----|------|-------|
39| 1 | SCOPE | Y | Y | Y | Y |
40| 2 | PLAN | - | Y | Y | Y |
41| 3 | RETRIEVE | Y | Y | Y | Y |
42| 4 | TRIANGULATE | - | Y | Y | Y |
43| 4.5 | OUTLINE REFINEMENT | - | Y | Y | Y |
44| 5 | SYNTHESIZE | - | Y | Y | Y |
45| 6 | CRITIQUE | - | - | Y | Y |
46| 7 | REFINE | - | - | Y | Y |
47| 8 | PACKAGE | Y | Y | Y | Y |
48
49**Note:** Phases 3-5 operate as an evidence loop per section (retrieve → evidence store → refine outline → draft → verify claims → delta-retrieve if needed), not as strict sequential gates.
50
51---
52
53## Execution
54
55**On invocation, load relevant reference files:**
56
571. **Phase 1-7:** Load [methodology.md](./reference/methodology.md) for detailed phase instructions
582. **Phase 8 (Report):** Load [report-assembly.md](./reference/report-assembly.md) for progressive generation
593. **HTML/PDF output:** Load [html-generation.md](./reference/html-generation.md)
604. **Quality checks:** Load [quality-gates.md](./reference/quality-gates.md)
615. **Long reports (>18K words):** Load [continuation.md](./reference/continuation.md)
62
63**Templates:**
64- Report structure: [report_template.md](./templates/report_template.md)
65- HTML styling: [mckinsey_report_template.html](./templates/mckinsey_report_template.html)
66
67**Scripts:**
68- `python scripts/validate_report.py --report [path]`
69- `python scripts/verify_citations.py --report [path]`
70- `python scripts/md_to_html.py [markdown_path]`
71
72---
73
74## Output Contract
75
76**Required sections:**
77- Executive Summary (200-400 words)
78- Introduction (scope, methodology, assumptions)
79- Main Analysis (4-8 findings, 600-2,000 words each, cited)
80- Synthesis & Insights (patterns, implications)
81- Limitations & Caveats
82- Recommendations
83- Bibliography (COMPLETE - every citation, no placeholders)
84- Methodology Appendix
85
86**Output files (all to `~/Documents/[Topic]_Research_[YYYYMMDD]/`):**
87- Markdown (primary source of truth)
88- `sources.jsonl` — stable source registry with canonical IDs
89- `evidence.jsonl` — append-only evidence store with quotes and locators
90- `claims.jsonl` — atomic claim ledger with support status
91- `run_manifest.json` — query, mode, assumptions, provider config
92- HTML (McKinsey style, auto-opened)
93- PDF (professional print, auto-opened)
94
95**Quality standards:**
96- 10+ sources, 3+ per major claim (cluster-independent, not just count)
97- All factual claims cited immediately [N] with evidence backing in `evidence.jsonl`
98- Claim-support verification mandatory: no unsupported factual claims pass delivery
99- No placeholders, no fabricated citations
100- Prose-first (>=80%), bullets sparingly
101
102---
103
104## When to Use / NOT Use
105
106**Use:** Comprehensive analysis, technology comparisons, state-of-the-art reviews, multi-perspective investigation, market analysis.
107
108**Do NOT use:** Simple lookups, debugging, 1-2 search answers, quick time-sensitive queries.