Microsoft GraphRAG
Overview
GraphRAG is a modular graph-based Retrieval-Augmented Generation (RAG) system developed by Microsoft Research. It extracts meaningful, structured data from unstructured text using LLMs to create a knowledge graph, then uses these connections to answer questions that span many documents or require thematic understanding. Unlike traditional vector-based RAG, GraphRAG excels at answering abstract questions like "What are the top themes in this dataset?" by leveraging community detection and hierarchical summarization.
Problem Addressed
| Problem |
Solution |
| Vector/keyword search fails on cross-document questions |
Knowledge graph connects information across large volumes of documents |
| Thematic questions unanswerable with traditional RAG |
Community detection and hierarchical summarization enable abstract reasoning |
| RAG limited to local context around retrieved chunks |
Global search via map-reduce over community reports provides dataset overview |
| Noisy data with conflicting information |
Graph structure surfaces relationships and identifies authoritative entities |
| Single retrieval strategy limits flexibility |
Multiple search modes: Local, Global, DRIFT, Basic for different query types |
| LLM costs high for re-indexing on errors |
Built-in LLM caching prevents redundant API calls during indexing |
| Lock-in to specific storage or model providers |
Factory pattern allows custom implementations for all subsystems |
Key Statistics
| Metric |
Value |
Date Gathered |
| GitHub Stars |
30,637 |
2026-01-31 |
| GitHub Forks |
3,229 |
2026-01-31 |
| Open Issues |
95 |
2026-01-31 |
| PyPI Monthly DL |
60,312 |
2026-01-31 |
| PyPI Weekly DL |
19,803 |
2026-01-31 |
| PyPI Daily DL |
4,051 |
2026-01-31 |
| Primary Language |
Python |
2026-01-31 |
| Repository Age |
Since March 2024 |
2026-01-31 |
| Python Required |
>=3.11, <3.14 |
2026-01-31 |
Key Features
Indexing Pipeline
- Entity extraction: LLM-powered extraction of entities, relationships, and claims from raw text
- Community detection: Graph-based clustering to identify related entity communities
- Hierarchical summarization: Community reports generated at multiple levels of granularity
- Chunk embedding: Text chunks embedded for vector similarity search
- Entity embedding: Entities embedded for semantic retrieval
- LLM caching: Cached completions for idempotent, resilient indexing
- Configurable workflows: Modular pipeline with customizable steps and prompts
Query Mechanisms
- Local Search: Combines AI-extracted knowledge graph with text chunks for entity-specific questions
- Global Search: Map-reduce over community reports for dataset-wide thematic questions
- DRIFT Search: Expands local search breadth using community insights for comprehensive answers
- Basic Search: Vector RAG baseline for comparison (top-k chunk retrieval)
- Question Generation: Generates follow-up questions for deeper investigation
Architecture
- Monorepo structure: Modular packages (graphrag-cache, graphrag-chunking, graphrag-common, graphrag-input, graphrag-llm, graphrag-storage, graphrag-vectors)
- Factory pattern: Extensible providers for models, storage, cache, vectors, input readers
- Parquet outputs: Indexes stored as Parquet tables for efficient querying
- CLI and Python API: Multiple interfaces for indexing and querying
Model Support
- OpenAI: Direct API integration
- Azure OpenAI: Full support with managed identity authentication
- LiteLLM wrapper: Extensible to additional providers
- Prompt tuning: Guidance for domain-specific prompt optimization
Storage Backends
- File storage: Local filesystem for development
- Azure Blob Storage: Cloud storage integration
- CosmosDB: Distributed database support
- Vector stores: LanceDB, Azure AI Search, CosmosDB vector search
Technical Architecture
Pipeline Flow
Load Documents
|
Chunk Documents
|
+-- Extract Graph --> Detect Communities --> Generate Reports --> Embed Reports
|
+-- Extract Claims
|
+-- Embed Chunks
|
+-- Embed Entities
Knowledge Model
| Component |
Description |
| Documents |
Source text files (txt, CSV, JSON) |
| Text Units |
Chunked document segments |
| Entities |
Extracted named entities (people, places, things) |
| Relationships |
Connections between entities |
| Claims |
Factual assertions extracted from text |
| Communities |
Clusters of related entities from graph analysis |
| Community Reports |
LLM-generated summaries at multiple hierarchy levels |
| Embeddings |
Vector representations for semantic search |
Factory Extensions
| Subsystem |
Purpose |
Built-in Options |
| Language Model |
Chat and embed methods |
LiteLLM wrapper (OpenAI, Azure, etc.) |
| Input Reader |
Document ingestion |
Text, CSV, JSON |
| Cache |
LLM response caching |
File, Blob, CosmosDB |
| Storage |
Index persistence |
File, Blob, CosmosDB |
| Vector Store |
Embedding storage and retrieval |
LanceDB, Azure AI Search, CosmosDB |
| Workflows |
Pipeline step customization |
Default GraphRAG pipeline |
Installation and Usage
Installation
# Using pip
pip install graphrag
# Using uv
uv pip install graphrag
Quick Start
# Initialize workspace
mkdir graphrag_project && cd graphrag_project
graphrag init
# Configure API key in .env
# GRAPHRAG_API_KEY=<your-api-key>
# Add documents to ./input directory
curl https://www.gutenberg.org/cache/epub/24022/pg24022.txt -o ./input/book.txt
# Run indexing
graphrag index
# Query with Global Search (thematic questions)
graphrag query "What are the top themes in this story?"
# Query with Local Search (entity-specific questions)
graphrag query "Who is Scrooge and what are his main relationships?" --method local
Python API
from graphrag.api import build_index, local_search, global_search
# Index documents
await build_index(root="./graphrag_project")
# Local search for entity questions
result = await local_search(
root="./graphrag_project",
query="What are the healing properties of chamomile?"
)
# Global search for thematic questions
result = await global_search(
root="./graphrag_project",
query="What are the main themes across all documents?"
)
Azure OpenAI Configuration
# settings.yaml
models:
default_chat_model:
type: chat
model_provider: azure
model: gpt-4.1
deployment_name: <AZURE_DEPLOYMENT_NAME>
api_base: https://<instance>.openai.azure.com
api_version: 2024-02-15-preview
auth_type: azure_managed_identity # Optional: for managed auth
Relevance to Claude Code Development
Direct Applications
Knowledge-Augmented Skills: GraphRAG patterns could inform how Claude Code skills organize and retrieve reference documentation, using graph structure instead of flat files.
Cross-Document Reasoning: The community detection and hierarchical summarization approach provides a model for answering questions that span multiple skill reference files.
Thematic Query Support: Global search mechanism could inspire how Claude Code answers abstract questions about a codebase ("What patterns does this project follow?").
Context Window Optimization: Community reports provide compressed summaries that could reduce token usage while maintaining comprehensive coverage.
Caching Patterns: LLM caching strategy applicable to Claude Code for expensive operations.
Patterns Worth Adopting
Multiple Search Strategies: Local vs Global vs DRIFT search demonstrates value of query-type-specific retrieval strategies.
Community Detection: Graph clustering to identify related concepts could improve skill organization and cross-referencing.
Hierarchical Summarization: Multi-level community reports enable both detailed and overview queries on same index.
Factory Pattern: Extensible providers for all subsystems enable customization without core changes.
Parquet Outputs: Columnar format for index storage enables efficient analytical queries.
Integration Opportunities
MCP Server: GraphRAG could be exposed as an MCP tool for Claude Code to query indexed knowledge bases.
Skill Indexing: Index skill reference documentation with GraphRAG for enhanced retrieval in complex multi-skill queries.
Codebase Understanding: Index codebase documentation/comments to answer architectural questions.
Research Aggregation: Index research entries (like this one) for cross-resource thematic queries.
Comparison: GraphRAG vs Traditional RAG
| Aspect |
GraphRAG |
Traditional Vector RAG |
| Cross-document reasoning |
Strong (graph connections) |
Weak (isolated chunks) |
| Thematic/abstract queries |
Strong (community reports) |
Weak (no global context) |
| Entity-specific queries |
Strong (local search + graph) |
Moderate (chunk retrieval) |
| Indexing cost |
High (LLM-intensive extraction) |
Low (embedding only) |
| Query latency |
Higher (map-reduce for global) |
Lower (single retrieval) |
| Storage requirements |
Higher (graph + reports + embeddings) |
Lower (embeddings only) |
| Update complexity |
Higher (re-extraction) |
Lower (re-embed changed docs) |
Cost Considerations
The README explicitly warns: "GraphRAG indexing can be an expensive operation." Best practices:
- Start with small test datasets
- Use inexpensive/fast models for initial experimentation
- Leverage LLM caching to avoid redundant API calls
- Consider update strategies before large indexing jobs
References
Research Method: Information gathered from official GitHub repository README, RAI transparency document, documentation pages (query overview, index overview, architecture), GitHub API for statistics, PyPI API for package info and download statistics. All claims verified against primary sources.
Freshness Tracking
| Field |
Value |
| Version Documented |
v3.0.1 |
| Release Date |
2026-01-28 |
| GitHub Stars |
30,637 (as of 2026-01-31) |
| Monthly Downloads |
60,312 (as of 2026-01-31) |
| Next Review Date |
2026-05-01 |
Review Triggers:
- Major version release (v4.x)
- Significant new query mechanisms
- New storage or model provider integrations
- GitHub stars milestone (40K, 50K)
- PyPI downloads milestone (100K monthly)
- New community detection algorithms
- Breaking changes to knowledge model or API
- Integration with Claude/Anthropic models
1---2name: microsoft-graphrag-23description: GraphRAG is a modular graph-based Retrieval-Augmented Generation (RAG) system developed by Microsoft Research.4---5# Microsoft GraphRAG67| Field | Value |8| ------------- | ----------------------------------------------------------------------------------------------------------- |9| Research Date | 2026-01-31 |10| Primary URL | <https://microsoft.github.io/graphrag/> |11| GitHub | <https://github.com/microsoft/graphrag> |12| PyPI | <https://pypi.org/project/graphrag/> |13| arXiv Paper | <https://arxiv.org/pdf/2404.16130> |14| Version | v3.0.1 (released 2026-01-28) |15| License | MIT |16| Blog Post | <https://www.microsoft.com/en-us/research/blog/graphrag-unlocking-llm-discovery-on-narrative-private-data/> |1718---1920## Overview2122GraphRAG is a modular graph-based Retrieval-Augmented Generation (RAG) system developed by Microsoft Research. It extracts meaningful, structured data from unstructured text using LLMs to create a knowledge graph, then uses these connections to answer questions that span many documents or require thematic understanding. Unlike traditional vector-based RAG, GraphRAG excels at answering abstract questions like "What are the top themes in this dataset?" by leveraging community detection and hierarchical summarization.2324---2526## Problem Addressed2728| Problem | Solution |29| ------------------------------------------------------- | ----------------------------------------------------------------------------- |30| Vector/keyword search fails on cross-document questions | Knowledge graph connects information across large volumes of documents |31| Thematic questions unanswerable with traditional RAG | Community detection and hierarchical summarization enable abstract reasoning |32| RAG limited to local context around retrieved chunks | Global search via map-reduce over community reports provides dataset overview |33| Noisy data with conflicting information | Graph structure surfaces relationships and identifies authoritative entities |34| Single retrieval strategy limits flexibility | Multiple search modes: Local, Global, DRIFT, Basic for different query types |35| LLM costs high for re-indexing on errors | Built-in LLM caching prevents redundant API calls during indexing |36| Lock-in to specific storage or model providers | Factory pattern allows custom implementations for all subsystems |3738---3940## Key Statistics4142| Metric | Value | Date Gathered |43| ---------------- | ---------------- | ------------- |44| GitHub Stars | 30,637 | 2026-01-31 |45| GitHub Forks | 3,229 | 2026-01-31 |46| Open Issues | 95 | 2026-01-31 |47| PyPI Monthly DL | 60,312 | 2026-01-31 |48| PyPI Weekly DL | 19,803 | 2026-01-31 |49| PyPI Daily DL | 4,051 | 2026-01-31 |50| Primary Language | Python | 2026-01-31 |51| Repository Age | Since March 2024 | 2026-01-31 |52| Python Required | >=3.11, <3.14 | 2026-01-31 |5354---5556## Key Features5758### Indexing Pipeline5960- **Entity extraction**: LLM-powered extraction of entities, relationships, and claims from raw text61- **Community detection**: Graph-based clustering to identify related entity communities62- **Hierarchical summarization**: Community reports generated at multiple levels of granularity63- **Chunk embedding**: Text chunks embedded for vector similarity search64- **Entity embedding**: Entities embedded for semantic retrieval65- **LLM caching**: Cached completions for idempotent, resilient indexing66- **Configurable workflows**: Modular pipeline with customizable steps and prompts6768### Query Mechanisms6970- **Local Search**: Combines AI-extracted knowledge graph with text chunks for entity-specific questions71- **Global Search**: Map-reduce over community reports for dataset-wide thematic questions72- **DRIFT Search**: Expands local search breadth using community insights for comprehensive answers73- **Basic Search**: Vector RAG baseline for comparison (top-k chunk retrieval)74- **Question Generation**: Generates follow-up questions for deeper investigation7576### Architecture7778- **Monorepo structure**: Modular packages (graphrag-cache, graphrag-chunking, graphrag-common, graphrag-input, graphrag-llm, graphrag-storage, graphrag-vectors)79- **Factory pattern**: Extensible providers for models, storage, cache, vectors, input readers80- **Parquet outputs**: Indexes stored as Parquet tables for efficient querying81- **CLI and Python API**: Multiple interfaces for indexing and querying8283### Model Support8485- **OpenAI**: Direct API integration86- **Azure OpenAI**: Full support with managed identity authentication87- **LiteLLM wrapper**: Extensible to additional providers88- **Prompt tuning**: Guidance for domain-specific prompt optimization8990### Storage Backends9192- **File storage**: Local filesystem for development93- **Azure Blob Storage**: Cloud storage integration94- **CosmosDB**: Distributed database support95- **Vector stores**: LanceDB, Azure AI Search, CosmosDB vector search9697---9899## Technical Architecture100101### Pipeline Flow102103```text104Load Documents105 |106Chunk Documents107 |108 +-- Extract Graph --> Detect Communities --> Generate Reports --> Embed Reports109 |110 +-- Extract Claims111 |112 +-- Embed Chunks113 |114 +-- Embed Entities115```116117### Knowledge Model118119| Component | Description |120| ----------------- | ---------------------------------------------------- |121| Documents | Source text files (txt, CSV, JSON) |122| Text Units | Chunked document segments |123| Entities | Extracted named entities (people, places, things) |124| Relationships | Connections between entities |125| Claims | Factual assertions extracted from text |126| Communities | Clusters of related entities from graph analysis |127| Community Reports | LLM-generated summaries at multiple hierarchy levels |128| Embeddings | Vector representations for semantic search |129130### Factory Extensions131132| Subsystem | Purpose | Built-in Options |133| -------------- | ------------------------------- | ------------------------------------- |134| Language Model | Chat and embed methods | LiteLLM wrapper (OpenAI, Azure, etc.) |135| Input Reader | Document ingestion | Text, CSV, JSON |136| Cache | LLM response caching | File, Blob, CosmosDB |137| Storage | Index persistence | File, Blob, CosmosDB |138| Vector Store | Embedding storage and retrieval | LanceDB, Azure AI Search, CosmosDB |139| Workflows | Pipeline step customization | Default GraphRAG pipeline |140141---142143## Installation and Usage144145### Installation146147```bash148# Using pip149pip install graphrag150151# Using uv152uv pip install graphrag153```154155### Quick Start156157```bash158# Initialize workspace159mkdir graphrag_project && cd graphrag_project160graphrag init161162# Configure API key in .env163# GRAPHRAG_API_KEY=<your-api-key>164165# Add documents to ./input directory166curl https://www.gutenberg.org/cache/epub/24022/pg24022.txt -o ./input/book.txt167168# Run indexing169graphrag index170171# Query with Global Search (thematic questions)172graphrag query "What are the top themes in this story?"173174# Query with Local Search (entity-specific questions)175graphrag query "Who is Scrooge and what are his main relationships?" --method local176```177178### Python API179180```python181from graphrag.api import build_index, local_search, global_search182183# Index documents184await build_index(root="./graphrag_project")185186# Local search for entity questions187result = await local_search(188 root="./graphrag_project",189 query="What are the healing properties of chamomile?"190)191192# Global search for thematic questions193result = await global_search(194 root="./graphrag_project",195 query="What are the main themes across all documents?"196)197```198199### Azure OpenAI Configuration200201```yaml202# settings.yaml203models:204 default_chat_model:205 type: chat206 model_provider: azure207 model: gpt-4.1208 deployment_name: <AZURE_DEPLOYMENT_NAME>209 api_base: https://<instance>.openai.azure.com210 api_version: 2024-02-15-preview211 auth_type: azure_managed_identity # Optional: for managed auth212```213214---215216## Relevance to Claude Code Development217218### Direct Applications2192201. **Knowledge-Augmented Skills**: GraphRAG patterns could inform how Claude Code skills organize and retrieve reference documentation, using graph structure instead of flat files.2212222. **Cross-Document Reasoning**: The community detection and hierarchical summarization approach provides a model for answering questions that span multiple skill reference files.2232243. **Thematic Query Support**: Global search mechanism could inspire how Claude Code answers abstract questions about a codebase ("What patterns does this project follow?").2252264. **Context Window Optimization**: Community reports provide compressed summaries that could reduce token usage while maintaining comprehensive coverage.2272285. **Caching Patterns**: LLM caching strategy applicable to Claude Code for expensive operations.229230### Patterns Worth Adopting2312321. **Multiple Search Strategies**: Local vs Global vs DRIFT search demonstrates value of query-type-specific retrieval strategies.2332342. **Community Detection**: Graph clustering to identify related concepts could improve skill organization and cross-referencing.2352363. **Hierarchical Summarization**: Multi-level community reports enable both detailed and overview queries on same index.2372384. **Factory Pattern**: Extensible providers for all subsystems enable customization without core changes.2392405. **Parquet Outputs**: Columnar format for index storage enables efficient analytical queries.241242### Integration Opportunities2432441. **MCP Server**: GraphRAG could be exposed as an MCP tool for Claude Code to query indexed knowledge bases.2452462. **Skill Indexing**: Index skill reference documentation with GraphRAG for enhanced retrieval in complex multi-skill queries.2472483. **Codebase Understanding**: Index codebase documentation/comments to answer architectural questions.2492504. **Research Aggregation**: Index research entries (like this one) for cross-resource thematic queries.251252### Comparison: GraphRAG vs Traditional RAG253254| Aspect | GraphRAG | Traditional Vector RAG |255| ------------------------- | ------------------------------------- | ----------------------------- |256| Cross-document reasoning | Strong (graph connections) | Weak (isolated chunks) |257| Thematic/abstract queries | Strong (community reports) | Weak (no global context) |258| Entity-specific queries | Strong (local search + graph) | Moderate (chunk retrieval) |259| Indexing cost | High (LLM-intensive extraction) | Low (embedding only) |260| Query latency | Higher (map-reduce for global) | Lower (single retrieval) |261| Storage requirements | Higher (graph + reports + embeddings) | Lower (embeddings only) |262| Update complexity | Higher (re-extraction) | Lower (re-embed changed docs) |263264### Cost Considerations265266The README explicitly warns: "GraphRAG indexing can be an expensive operation." Best practices:267268- Start with small test datasets269- Use inexpensive/fast models for initial experimentation270- Leverage LLM caching to avoid redundant API calls271- Consider update strategies before large indexing jobs272273---274275## References276277| Source | URL | Accessed |278| -------------------------- | ----------------------------------------------------------------------------------------------------------- | ---------- |279| GitHub Repository | <https://github.com/microsoft/graphrag> | 2026-01-31 |280| Official Documentation | <https://microsoft.github.io/graphrag/> | 2026-01-31 |281| arXiv Paper | <https://arxiv.org/pdf/2404.16130> | 2026-01-31 |282| Microsoft Research Blog | <https://www.microsoft.com/en-us/research/blog/graphrag-unlocking-llm-discovery-on-narrative-private-data/> | 2026-01-31 |283| PyPI Package | <https://pypi.org/project/graphrag/> | 2026-01-31 |284| PyPI Stats | <https://pypistats.org/packages/graphrag> | 2026-01-31 |285| RAI Transparency Document | <https://github.com/microsoft/graphrag/blob/main/RAI_TRANSPARENCY.md> | 2026-01-31 |286| Query Engine Documentation | <https://microsoft.github.io/graphrag/query/overview/> | 2026-01-31 |287| Indexing Documentation | <https://microsoft.github.io/graphrag/index/overview/> | 2026-01-31 |288| Architecture Documentation | <https://microsoft.github.io/graphrag/index/architecture/> | 2026-01-31 |289290**Research Method**: Information gathered from official GitHub repository README, RAI transparency document, documentation pages (query overview, index overview, architecture), GitHub API for statistics, PyPI API for package info and download statistics. All claims verified against primary sources.291292---293294## Freshness Tracking295296| Field | Value |297| ------------------ | ------------------------- |298| Version Documented | v3.0.1 |299| Release Date | 2026-01-28 |300| GitHub Stars | 30,637 (as of 2026-01-31) |301| Monthly Downloads | 60,312 (as of 2026-01-31) |302| Next Review Date | 2026-05-01 |303304**Review Triggers**:305306- Major version release (v4.x)307- Significant new query mechanisms308- New storage or model provider integrations309- GitHub stars milestone (40K, 50K)310- PyPI downloads milestone (100K monthly)311- New community detection algorithms312- Breaking changes to knowledge model or API313- Integration with Claude/Anthropic models