LlamaIndex
Overview
LlamaIndex is a data framework for building RAG pipelines, knowledge assistants, and data-augmented LLM applications. It provides document loading from 300+ sources, flexible chunking strategies, multiple index types, hybrid retrieval with reranking, and production evaluation tools for question-answering systems.
Instructions
- When ingesting documents, use
SimpleDirectoryReader for local files or LlamaHub connectors for SaaS platforms, and run through an IngestionPipeline with metadata extractors (title, summary) and deduplication.
- When chunking, start with
SentenceSplitter at 1024 tokens with 200 token overlap, use MarkdownNodeParser for structured documents, CodeSplitter for code, and adjust based on evaluation results.
- When indexing, use
VectorStoreIndex as the default for most RAG, KnowledgeGraphIndex for entity relationships, and DocumentSummaryIndex for per-document summaries.
- When retrieving, implement hybrid retrieval (vector + keyword) for production, add a reranker (
CohereRerank) after retrieval for improved relevance, and set similarity_top_k based on context window (3-5 for large models, 2-3 for smaller).
- When building query engines, use
RetrieverQueryEngine for standard RAG, CitationQueryEngine for responses with source attribution, and SubQuestionQueryEngine for complex multi-part queries.
- When creating agents, use
ReActAgent with tools wrapping query engines (QueryEngineTool), functions, and other agents for multi-step reasoning.
- When evaluating, use
CorrectnessEvaluator, FaithfulnessEvaluator, and RelevancyEvaluator on a test set before deploying.
Examples
Example 1: Build a RAG pipeline over company documentation
User request: "Create a question-answering system over our internal docs"
Actions:
- Load documents with
SimpleDirectoryReader and extract metadata (title, summary)
- Chunk with
SentenceSplitter (1024 tokens, 200 overlap) through an IngestionPipeline
- Create
VectorStoreIndex with OpenAI embeddings and configure hybrid retrieval
- Build
CitationQueryEngine for answers with source references
Output: A RAG system that answers questions with citations from company documentation.
Example 2: Create a multi-source research agent
User request: "Build an agent that can search across our docs, database, and web"
Actions:
- Create separate query engines for each data source (vector index, SQL, web search)
- Wrap each engine as a
QueryEngineTool with descriptive tool descriptions
- Build a
ReActAgent that routes questions to the appropriate tool
- Add
SubQuestionQueryEngine for complex queries requiring multiple sources
Output: An intelligent agent that reasons about which data source to query and synthesizes multi-source answers.
Guidelines
- Use
SentenceSplitter with 1024 token chunks and 200 token overlap as the starting point.
- Always add metadata extractors to the ingestion pipeline; title and summary metadata improve retrieval significantly.
- Use hybrid retrieval (vector + keyword) for production; pure vector search misses exact term matches.
- Add a reranker (
CohereRerank) after retrieval to improve result relevance for small cost.
- Evaluate with
CorrectnessEvaluator on a test set before deploying; subjective quality assessment does not scale.
- Set
similarity_top_k based on context window: 3-5 chunks for large models, 2-3 for smaller models.
- Use
IngestionPipeline with deduplication for incremental data updates; do not re-embed unchanged documents.
1---2name: llamaindex3description: LlamaIndex4---5# LlamaIndex67## Overview89LlamaIndex is a data framework for building RAG pipelines, knowledge assistants, and data-augmented LLM applications. It provides document loading from 300+ sources, flexible chunking strategies, multiple index types, hybrid retrieval with reranking, and production evaluation tools for question-answering systems.1011## Instructions1213- When ingesting documents, use `SimpleDirectoryReader` for local files or LlamaHub connectors for SaaS platforms, and run through an `IngestionPipeline` with metadata extractors (title, summary) and deduplication.14- When chunking, start with `SentenceSplitter` at 1024 tokens with 200 token overlap, use `MarkdownNodeParser` for structured documents, `CodeSplitter` for code, and adjust based on evaluation results.15- When indexing, use `VectorStoreIndex` as the default for most RAG, `KnowledgeGraphIndex` for entity relationships, and `DocumentSummaryIndex` for per-document summaries.16- When retrieving, implement hybrid retrieval (vector + keyword) for production, add a reranker (`CohereRerank`) after retrieval for improved relevance, and set `similarity_top_k` based on context window (3-5 for large models, 2-3 for smaller).17- When building query engines, use `RetrieverQueryEngine` for standard RAG, `CitationQueryEngine` for responses with source attribution, and `SubQuestionQueryEngine` for complex multi-part queries.18- When creating agents, use `ReActAgent` with tools wrapping query engines (`QueryEngineTool`), functions, and other agents for multi-step reasoning.19- When evaluating, use `CorrectnessEvaluator`, `FaithfulnessEvaluator`, and `RelevancyEvaluator` on a test set before deploying.2021## Examples2223### Example 1: Build a RAG pipeline over company documentation2425**User request:** "Create a question-answering system over our internal docs"2627**Actions:**281. Load documents with `SimpleDirectoryReader` and extract metadata (title, summary)292. Chunk with `SentenceSplitter` (1024 tokens, 200 overlap) through an `IngestionPipeline`303. Create `VectorStoreIndex` with OpenAI embeddings and configure hybrid retrieval314. Build `CitationQueryEngine` for answers with source references3233**Output:** A RAG system that answers questions with citations from company documentation.3435### Example 2: Create a multi-source research agent3637**User request:** "Build an agent that can search across our docs, database, and web"3839**Actions:**401. Create separate query engines for each data source (vector index, SQL, web search)412. Wrap each engine as a `QueryEngineTool` with descriptive tool descriptions423. Build a `ReActAgent` that routes questions to the appropriate tool434. Add `SubQuestionQueryEngine` for complex queries requiring multiple sources4445**Output:** An intelligent agent that reasons about which data source to query and synthesizes multi-source answers.4647## Guidelines4849- Use `SentenceSplitter` with 1024 token chunks and 200 token overlap as the starting point.50- Always add metadata extractors to the ingestion pipeline; title and summary metadata improve retrieval significantly.51- Use hybrid retrieval (vector + keyword) for production; pure vector search misses exact term matches.52- Add a reranker (`CohereRerank`) after retrieval to improve result relevance for small cost.53- Evaluate with `CorrectnessEvaluator` on a test set before deploying; subjective quality assessment does not scale.54- Set `similarity_top_k` based on context window: 3-5 chunks for large models, 2-3 for smaller models.55- Use `IngestionPipeline` with deduplication for incremental data updates; do not re-embed unchanged documents.