title: "Haystack"
description: "Production-ready NLP framework for building search pipelines, RAG systems, and LLM applications. Modular pipeline architecture with 100+ integrations (OpenAI, Cohere, Weaviate, Elasticsearch, Pinecone). Best for enterprise search, document Q&A, and complex multi-step RAG pipelines."
skillName: "haystack"
skillVersion: "1.0.0"
skillAuthor: "Orchestra Research"
skillLicense: "MIT"
skillTags: ["RAG", "Haystack", "NLP", "Search Pipelines", "Document QA", "Enterprise Search", "Pipeline Architecture", "Deepset", "Production"]
skillDeps: ["haystack-ai", "transformers", "sentence-transformers"]
|
|
| Version |
1.0.0 |
| Author |
Orchestra Research |
| License |
MIT |
| Tags |
RAG Haystack NLP Search Pipelines Document QA Enterprise Search |
| Dependencies |
haystack-ai transformers sentence-transformers |
Haystack — Production NLP & RAG Pipelines
deepset's modular framework for building search systems and RAG applications at production scale.
When to use Haystack
Use Haystack when:
- Building production-grade search or Q&A systems
- Complex multi-step pipelines (retrieve → rerank → generate)
- Enterprise document Q&A over large corpora
- Need extensive integrations (100+ document stores, LLMs, embedders)
- Evaluation-driven RAG development
Metrics:
- 18,000+ GitHub stars
- 100+ integrations (OpenAI, Cohere, Hugging Face, Elasticsearch, Weaviate, Pinecone)
- Production-battle-tested at enterprise scale
- Apache 2.0 license
Quick start
pip install haystack-ai
# With OpenAI
pip install haystack-ai openai
# With local models
pip install haystack-ai transformers sentence-transformers
Basic RAG pipeline
from haystack import Pipeline, Document
from haystack.document_stores.in_memory import InMemoryDocumentStore
from haystack.components.retrievers import InMemoryBM25Retriever
from haystack.components.generators import OpenAIGenerator
from haystack.components.builders import PromptBuilder
# 1. Set up document store
store = InMemoryDocumentStore()
store.write_documents([
Document(content="Flash Attention reduces memory from O(N²) to O(N) by using tiling and recomputation."),
Document(content="GRPO uses group-relative policy optimization for RL training of LLMs."),
Document(content="LoRA adds low-rank matrices to frozen weights, reducing trainable parameters by 10,000x."),
])
# 2. Build pipeline
template = """Answer the question based on the context.
Context: {% for doc in documents %}{{ doc.content }}{% endfor %}
Question: {{ question }}"""
pipeline = Pipeline()
pipeline.add_component("retriever", InMemoryBM25Retriever(document_store=store))
pipeline.add_component("prompt_builder", PromptBuilder(template=template))
pipeline.add_component("llm", OpenAIGenerator(model="gpt-4o-mini"))
pipeline.connect("retriever", "prompt_builder.documents")
pipeline.connect("prompt_builder", "llm")
# 3. Run
result = pipeline.run({
"retriever": {"query": "How does Flash Attention save memory?"},
"prompt_builder": {"question": "How does Flash Attention save memory?"}
})
print(result["llm"]["replies"][0])
Semantic search with embeddings
from haystack.components.embedders import SentenceTransformersDocumentEmbedder, SentenceTransformersTextEmbedder
from haystack.components.retrievers.in_memory import InMemoryEmbeddingRetriever
from haystack.document_stores.in_memory import InMemoryDocumentStore
# Indexing pipeline
indexing = Pipeline()
indexing.add_component("embedder", SentenceTransformersDocumentEmbedder(
model="sentence-transformers/all-MiniLM-L6-v2"
))
indexing.add_component("writer", DocumentWriter(document_store=store))
indexing.connect("embedder", "writer")
indexing.run({"embedder": {"documents": documents}})
# Query pipeline
query_pipeline = Pipeline()
query_pipeline.add_component("text_embedder", SentenceTransformersTextEmbedder(
model="sentence-transformers/all-MiniLM-L6-v2"
))
query_pipeline.add_component("retriever", InMemoryEmbeddingRetriever(document_store=store, top_k=5))
query_pipeline.connect("text_embedder.embedding", "retriever.query_embedding")
results = query_pipeline.run({"text_embedder": {"text": "attention mechanism memory optimization"}})
Advanced: Hybrid retrieval + reranking
from haystack.components.rankers import TransformersSimilarityRanker
from haystack.components.joiners import DocumentJoiner
pipeline = Pipeline()
pipeline.add_component("bm25_retriever", InMemoryBM25Retriever(document_store=store, top_k=10))
pipeline.add_component("embedding_retriever", InMemoryEmbeddingRetriever(document_store=store, top_k=10))
pipeline.add_component("joiner", DocumentJoiner(join_mode="reciprocal_rank_fusion"))
pipeline.add_component("reranker", TransformersSimilarityRanker(model="cross-encoder/ms-marco-MiniLM-L-6-v2", top_k=5))
pipeline.add_component("llm", OpenAIGenerator(model="gpt-4o"))
pipeline.connect("bm25_retriever", "joiner")
pipeline.connect("embedding_retriever", "joiner")
pipeline.connect("joiner", "reranker")
# ... connect to prompt builder and llm
Pipeline evaluation
from haystack.evaluation import EvaluationRunResult
from haystack.components.evaluators import ContextRelevanceEvaluator, FaithfulnessEvaluator
evaluator = Pipeline()
evaluator.add_component("context_relevance", ContextRelevanceEvaluator(api="openai"))
evaluator.add_component("faithfulness", FaithfulnessEvaluator(api="openai"))
results = evaluator.run({
"context_relevance": {"questions": questions, "contexts": retrieved_docs},
"faithfulness": {"questions": questions, "contexts": retrieved_docs, "responses": answers}
})
Serialization & deployment
import yaml
# Save pipeline as YAML
with open("pipeline.yaml", "w") as f:
pipeline.dump(f)
# Load and run
with open("pipeline.yaml") as f:
pipeline = Pipeline.load(f)
# REST API with Hayhooks
# pip install hayhooks
# hayhooks run pipeline.yaml
Common pitfalls
- Slow first run: Embedding models download on first use (~500MB); cache with
HF_HOME
- Pipeline connection errors: Check component input/output socket names exactly
- BM25 vs semantic: BM25 better for keyword search; embeddings better for conceptual
- Memory on large corpora: Use Elasticsearch/Weaviate instead of InMemoryDocumentStore
References
1---2name: haystack3description: ---4---5---6 title: "Haystack"7 description: "Production-ready NLP framework for building search pipelines, RAG systems, and LLM applications. Modular pipeline architecture with 100+ integrations (OpenAI, Cohere, Weaviate, Elasticsearch, Pinecone). Best for enterprise search, document Q&A, and complex multi-step RAG pipelines."8 skillName: "haystack"9 skillVersion: "1.0.0"10 skillAuthor: "Orchestra Research"11 skillLicense: "MIT"12 skillTags: ["RAG", "Haystack", "NLP", "Search Pipelines", "Document QA", "Enterprise Search", "Pipeline Architecture", "Deepset", "Production"]13 skillDeps: ["haystack-ai", "transformers", "sentence-transformers"]14 ---1516 | | |17 |---|---|18 | **Version** | 1.0.0 |19 | **Author** | Orchestra Research |20 | **License** | MIT |21 | **Tags** | `RAG` `Haystack` `NLP` `Search Pipelines` `Document QA` `Enterprise Search` |22 | **Dependencies** | `haystack-ai` `transformers` `sentence-transformers` |232425 # Haystack — Production NLP & RAG Pipelines2627 deepset's modular framework for building search systems and RAG applications at production scale.2829 ## When to use Haystack3031 **Use Haystack when:**32 - Building production-grade search or Q&A systems33 - Complex multi-step pipelines (retrieve → rerank → generate)34 - Enterprise document Q&A over large corpora35 - Need extensive integrations (100+ document stores, LLMs, embedders)36 - Evaluation-driven RAG development3738 **Metrics**:39 - **18,000+ GitHub stars**40 - **100+ integrations** (OpenAI, Cohere, Hugging Face, Elasticsearch, Weaviate, Pinecone)41 - **Production-battle-tested** at enterprise scale42 - Apache 2.0 license4344 ## Quick start4546 ```bash47 pip install haystack-ai48 # With OpenAI49 pip install haystack-ai openai50 # With local models51 pip install haystack-ai transformers sentence-transformers52 ```5354 ### Basic RAG pipeline5556 ```python57 from haystack import Pipeline, Document58 from haystack.document_stores.in_memory import InMemoryDocumentStore59 from haystack.components.retrievers import InMemoryBM25Retriever60 from haystack.components.generators import OpenAIGenerator61 from haystack.components.builders import PromptBuilder6263 # 1. Set up document store64 store = InMemoryDocumentStore()65 store.write_documents([66 Document(content="Flash Attention reduces memory from O(N²) to O(N) by using tiling and recomputation."),67 Document(content="GRPO uses group-relative policy optimization for RL training of LLMs."),68 Document(content="LoRA adds low-rank matrices to frozen weights, reducing trainable parameters by 10,000x."),69 ])7071 # 2. Build pipeline72 template = """Answer the question based on the context.73 Context: {% for doc in documents %}{{ doc.content }}{% endfor %}74 Question: {{ question }}"""7576 pipeline = Pipeline()77 pipeline.add_component("retriever", InMemoryBM25Retriever(document_store=store))78 pipeline.add_component("prompt_builder", PromptBuilder(template=template))79 pipeline.add_component("llm", OpenAIGenerator(model="gpt-4o-mini"))8081 pipeline.connect("retriever", "prompt_builder.documents")82 pipeline.connect("prompt_builder", "llm")8384 # 3. Run85 result = pipeline.run({86 "retriever": {"query": "How does Flash Attention save memory?"},87 "prompt_builder": {"question": "How does Flash Attention save memory?"}88 })89 print(result["llm"]["replies"][0])90 ```9192 ### Semantic search with embeddings9394 ```python95 from haystack.components.embedders import SentenceTransformersDocumentEmbedder, SentenceTransformersTextEmbedder96 from haystack.components.retrievers.in_memory import InMemoryEmbeddingRetriever97 from haystack.document_stores.in_memory import InMemoryDocumentStore9899 # Indexing pipeline100 indexing = Pipeline()101 indexing.add_component("embedder", SentenceTransformersDocumentEmbedder(102 model="sentence-transformers/all-MiniLM-L6-v2"103 ))104 indexing.add_component("writer", DocumentWriter(document_store=store))105 indexing.connect("embedder", "writer")106107 indexing.run({"embedder": {"documents": documents}})108109 # Query pipeline110 query_pipeline = Pipeline()111 query_pipeline.add_component("text_embedder", SentenceTransformersTextEmbedder(112 model="sentence-transformers/all-MiniLM-L6-v2"113 ))114 query_pipeline.add_component("retriever", InMemoryEmbeddingRetriever(document_store=store, top_k=5))115 query_pipeline.connect("text_embedder.embedding", "retriever.query_embedding")116117 results = query_pipeline.run({"text_embedder": {"text": "attention mechanism memory optimization"}})118 ```119120 ### Advanced: Hybrid retrieval + reranking121122 ```python123 from haystack.components.rankers import TransformersSimilarityRanker124 from haystack.components.joiners import DocumentJoiner125126 pipeline = Pipeline()127 pipeline.add_component("bm25_retriever", InMemoryBM25Retriever(document_store=store, top_k=10))128 pipeline.add_component("embedding_retriever", InMemoryEmbeddingRetriever(document_store=store, top_k=10))129 pipeline.add_component("joiner", DocumentJoiner(join_mode="reciprocal_rank_fusion"))130 pipeline.add_component("reranker", TransformersSimilarityRanker(model="cross-encoder/ms-marco-MiniLM-L-6-v2", top_k=5))131 pipeline.add_component("llm", OpenAIGenerator(model="gpt-4o"))132133 pipeline.connect("bm25_retriever", "joiner")134 pipeline.connect("embedding_retriever", "joiner")135 pipeline.connect("joiner", "reranker")136 # ... connect to prompt builder and llm137 ```138139 ### Pipeline evaluation140141 ```python142 from haystack.evaluation import EvaluationRunResult143 from haystack.components.evaluators import ContextRelevanceEvaluator, FaithfulnessEvaluator144145 evaluator = Pipeline()146 evaluator.add_component("context_relevance", ContextRelevanceEvaluator(api="openai"))147 evaluator.add_component("faithfulness", FaithfulnessEvaluator(api="openai"))148149 results = evaluator.run({150 "context_relevance": {"questions": questions, "contexts": retrieved_docs},151 "faithfulness": {"questions": questions, "contexts": retrieved_docs, "responses": answers}152 })153 ```154155 ## Serialization & deployment156157 ```python158 import yaml159160 # Save pipeline as YAML161 with open("pipeline.yaml", "w") as f:162 pipeline.dump(f)163164 # Load and run165 with open("pipeline.yaml") as f:166 pipeline = Pipeline.load(f)167168 # REST API with Hayhooks169 # pip install hayhooks170 # hayhooks run pipeline.yaml171 ```172173 ## Common pitfalls174175 - **Slow first run**: Embedding models download on first use (~500MB); cache with `HF_HOME`176 - **Pipeline connection errors**: Check component input/output socket names exactly177 - **BM25 vs semantic**: BM25 better for keyword search; embeddings better for conceptual178 - **Memory on large corpora**: Use Elasticsearch/Weaviate instead of InMemoryDocumentStore179180 ## References181 - [Haystack GitHub](https://github.com/deepset-ai/haystack)182 - [Documentation](https://docs.haystack.deepset.ai/)183 - [Integrations Hub](https://haystack.deepset.ai/integrations)184 - [Tutorials](https://haystack.deepset.ai/tutorials)185