Skip to content
BotServBotServ
HaystackPipelineRetrieverGeneratorRAGPython

Haystack: End-to-End NLP Framework

Build search and RAG pipelines with Haystack. Learn retrievers, readers, generators, and run locally.

S

schutzgeist

7 min read
Haystack: End-to-End NLP Framework

Haystack: End-to-End NLP Framework

What This Article Covers

  • What Haystack is and how pipelines work
  • The roles of Document Store, Retriever, PromptBuilder, and Generator
  • Loading documents into an InMemoryDocumentStore
  • Building a RAG pipeline with Ollama
  • Common pitfalls and gotchas when getting started with Haystack 2.x

Introduction

Haystack is an open-source Python framework for building Natural Language Processing pipelines. It specializes in search applications, question answering, and RAG. Instead of writing glue code, you connect pre-built components in a pipeline. Each component has inputs and outputs that you wire together through defined connections.

Haystack 2.x simplified the architecture considerably. The central idea is straightforward: store documents, find relevant ones, build a prompt, and let a language model respond. The framework provides components like InMemoryDocumentStore, InMemoryBM25Retriever, PromptBuilder, and OllamaGenerator to handle this flow.

Why Use Haystack?

Once you’re combining multiple steps, a bare API approach quickly becomes hard to follow. Haystack makes the process visible. You can see how data flows through the pipeline, which components are involved, and where to make changes to improve results.

The modular design makes experimentation straightforward. You can swap out a retriever, test a different embedding model, or use another generator without rewriting your entire application. Haystack also integrates beautifully with Ollama, letting you avoid costs and data privacy concerns.

Haystack at a Glance

Haystack centers around a few core concepts:

  • Document: A text object with content and optional metadata.
  • Document Store: Storage for Document objects.
  • Component: A pipeline building block with defined inputs and outputs.
  • Pipeline: A directed graph of connected components.
  • Retriever: Finds relevant documents from a Document Store.
  • Reader: A model that extracts answers from documents.
  • PromptBuilder: Constructs a prompt from templates and variables.
  • Generator: A language model that produces text.
  • Embedder: Computes embeddings for documents or queries.

The pipeline is the heart of everything. You add components, connect their outputs, and run the workflow with pipeline.run().

Who This Article Is For

This article assumes you have basic Python knowledge. You don’t need to be an NLP expert, but you should be comfortable installing packages and running simple scripts. If you’ve worked with LangChain or LlamaIndex, the modular design will feel familiar.

Beginners benefit from Haystack’s way of breaking processes into visible steps. If you want to dive deeper into RAG, you’ll find more details in RAG Fundamentals and Local RAG.

Key Terms

TermDefinition
DocumentA text block with metadata stored in the Document Store.
Document StoreStorage that manages and searches many documents.
ComponentA single pipeline building block with inputs and outputs.
PipelineA collection of components that processes data.
RetrieverComponent that finds documents matching a query.
PromptBuilderConstructs a prompt from a template and variables.
GeneratorAn LLM that generates text from a prompt.
EmbedderComputes vectors for documents or queries.
ReaderA model that reads documents and extracts answers.

Installation

For Haystack 2.x with Ollama, you need these packages:

pip install haystack-ai ollama-haystack

For PDF support, add:

pip install pypdf

Make sure Ollama is running and your desired model is available. For example:

ollama pull llama3.2
ollama pull nomic-embed-text

Loading Documents into the Store

Before building a RAG pipeline, you need to load documents into a Document Store. InMemoryDocumentStore is perfectly fine for getting started.

from haystack import Document
from haystack.document_stores.in_memory import InMemoryDocumentStore

document_store = InMemoryDocumentStore()

dokumente = [
    Document(content="LlamaIndex focuses strongly on RAG and indexing."),
    Document(content="Haystack is a modular NLP framework for pipelines."),
    Document(content="LangChain offers chains, tools, and agents."),
]

document_store.write_documents(dokumente)
print(f"Number of documents: {document_store.count_documents()}")

For production use, you can swap InMemoryDocumentStore for Elasticsearch, OpenSearch, Weaviate, or others. For local testing, in-memory storage works fine.

A Simple Search Pipeline

With a retriever, you can search documents without invoking a language model. BM25 is a classic retriever based on word matching.

from haystack import Pipeline
from haystack.components.retrievers.in_memory import InMemoryBM25Retriever

retriever = InMemoryBM25Retriever(document_store=document_store)

pipeline = Pipeline()
pipeline.add_component("retriever", retriever)

ergebnis = pipeline.run({"retriever": {"query": "Was ist Haystack?"}})
for doc in ergebnis["retriever"]["documents"]:
    print(doc.content)

This example shows the pipeline pattern: add a component, pass parameters, retrieve results. For semantic search, you’d swap the retriever for an embedding-based one later.

RAG Pipeline with Ollama

A complete RAG pipeline connects retrieval, prompt building, and generation. PromptBuilder uses a Jinja template where retrieved documents and the question are inserted.

from haystack import Pipeline
from haystack.components.builders.prompt_builder import PromptBuilder
from haystack.components.retrievers.in_memory import InMemoryBM25Retriever
from haystack_integrations.components.generators.ollama import OllamaGenerator

template = """
Answer the question solely based on the context below.

Context:
{% for document in documents %}
{{ document.content }}
{% endfor %}

Question: {{ query }}
Answer:
"""

document_store = InMemoryDocumentStore()
document_store.write_documents(dokumente)

pipeline = Pipeline()
pipeline.add_component("retriever", InMemoryBM25Retriever(document_store=document_store))
pipeline.add_component("prompt_builder", PromptBuilder(template=template))
pipeline.add_component("llm", OllamaGenerator(model="llama3.2", url="http://localhost:11434"))

pipeline.connect("retriever", "prompt_builder.documents")
pipeline.connect("prompt_builder", "llm")

frage = "Was macht Haystack besonders?"
antwort = pipeline.run({"retriever": {"query": frage}, "prompt_builder": {"query": frage}})

print(antwort["llm"]["replies"][0])

This pipeline executes four steps: retrieve matching documents, assemble the prompt, call the local model, and output the answer. You can customize the template to control how the model behaves.

Working with Embeddings

For semantic search, compute embeddings on your documents and use an embedding retriever. The OllamaDocumentEmbedder and OllamaTextEmbedder components handle this.

from haystack import Document, Pipeline
from haystack.document_stores.in_memory import InMemoryDocumentStore
from haystack_integrations.components.embedders.ollama import OllamaDocumentEmbedder, OllamaTextEmbedder
from haystack.components.retrievers.in_memory import InMemoryEmbeddingRetriever

document_store = InMemoryDocumentStore()

embedder = OllamaDocumentEmbedder(model="nomic-embed-text")

dokumente = [
    Document(content="Python ist eine vielseitige Programmiersprache."),
    Document(content="JavaScript wird häufig für Webanwendungen genutzt."),
]

dokumente_mit_embeddings = embedder.run(documents=dokumente)["documents"]
document_store.write_documents(dokumente_mit_embeddings)

query_embedder = OllamaTextEmbedder(model="nomic-embed-text")
retriever = InMemoryEmbeddingRetriever(document_store=document_store)

pipeline = Pipeline()
pipeline.add_component("query_embedder", query_embedder)
pipeline.add_component("retriever", retriever)
pipeline.connect("query_embedder", "retriever.query_embedding")

ergebnis = pipeline.run({"query_embedder": {"text": "Wozu eignet sich Python?"}})
for doc in ergebnis["retriever"]["documents"]:
    print(doc.content)

This approach retrieves documents with similar meaning, even when the query uses different wording than the stored content. Learn more in Embedding Models and Vector Databases.

Understanding Pipeline Connections

Haystack connections follow a simple pattern: pipeline.connect("component_a", "component_b.input"). This wires the output of A to a specific input of B. Missing connections cause errors, so it helps to inspect the pipeline graph before running it.

You can serialize and reuse pipelines too. pipeline.dumps() produces a YAML representation, letting you save, version, and share finished pipelines.

Common Gotchas

  • Wrong imports: Haystack 2.x uses different paths than version 1.x. Watch for haystack instead of old module names.
  • Missing connections: Every component must be connected to the right input with pipeline.connect.
  • Wrong input names: prompt_builder.documents and prompt_builder.prompt are different. Check the docs.
  • Template errors: Jinja templates must be valid. A missing {% endfor %} causes a syntax error.
  • Model unavailable: Make sure Ollama is running and the chosen model has been pulled.
  • Wrong URL: Ollama’s default port is 11434. url should point to http://localhost:11434.
  • Documents not written: An empty document store returns no results. Check document_store.count_documents().
  • Timeout too low: Local models need time. Increase timeouts if requests abort.

Hardware, Costs, and Security

Haystack itself is free. Costs come from infrastructure or Cloud APIs. Running locally, you pay for power and hardware. A 7B model typically runs on 8 GB VRAM with quantization. For small test sets, CPU mode works but slower.

Security starts with the document store. When indexing sensitive content, enforce access controls and encrypted storage. Avoid publicly exposed Ollama instances. Run anything handling confidential data in a protected network.

Further Resources

FAQ

Is Haystack free?

Yes, the framework is open source. Costs only arise from models, infrastructure, or Cloud APIs.

What’s the difference between Haystack 1.x and 2.x?

Haystack 2.x was rebuilt from scratch around components and pipelines. Many old import paths no longer work.

Can I run Haystack locally?

Yes. With Ollama and the ollama-haystack package, all calls stay local.

What retrievers are available?

Haystack offers BM25 and embedding retrievers, plus specialized ones for Elasticsearch, OpenSearch, and other stores.

What is a PromptBuilder?

The PromptBuilder injects variables like found documents and questions into a template, creating the final prompt.

Do I need Python knowledge?

Yes, most examples assume Python. Basics are enough to get started.

Can I save pipelines?

Yes. Use pipeline.dumps() to get a YAML representation you can load later.

What document formats can Haystack handle?

Via converters and preprocessors, you can ingest PDFs, text files, Word documents, and many other formats.

What’s the advantage over LangChain?

Haystack focuses on visible pipelines and retriever-reader chains. LangChain offers more freedom for agents and chains.

How can I improve my results?

Experiment with different retrievers, splitter settings, prompt templates, and models. Embedding quality often makes the difference.

Is there a graphical interface?

Haystack itself is text-based but integrates well with Streamlit, FastAPI, or Jupyter Notebooks.

What happens if no matching document is found?

The prompt stays empty or lacks relevant context. The model then responds from its training knowledge, which often hallucinates with local models. Always verify result quality.

Sources

Back to Blog
Share:

Related Posts