Skip to content
BotServBotServ
RAGTutorialGuideChromaOllamaPythonLocal AIPractical

RAG Tutorials: Step-by-Step Guides

Practical RAG tutorials: build your own knowledge base, make PDFs searchable, chat with documents. Complete guides with code.

S

schutzgeist

12 min read
RAG Tutorials: Step-by-Step Guides

RAG Tutorials: Step-by-Step Guides

What this article covers

  • Four complete tutorials ranging from minimal to interactive, all with runnable Python code
  • Setup instructions for Python, Ollama, and required packages
  • Practical tips for adapting tutorials to your own documents
  • Common pitfalls and how to avoid them
  • FAQ with frequent questions about RAG tutorials

Introduction: Understanding RAG tutorials

The theory behind RAG is straightforward: search documents, find relevant passages, pass them to a language model. But there’s often a gap between theory and working code. That’s where these RAG tutorials come in. You’ll get four complete walkthroughs that you can follow step by step. Each tutorial builds on the previous one, bringing you closer to a real RAG application.

If you’re new to the basics, start with the RAG Fundamentals article first. There I explain the concepts behind Retrieval Augmented Generation. If you’re wondering what local AI means in general, you’ll find answers in What is local AI?.

Why do you need tutorials?

Imagine you’ve read the RAG fundamentals. You understand embeddings, how vector databases work, and why chunking matters. Then you open your editor and ask yourself: where do I start? Which library do I need? How do I connect Ollama with Chroma? How do I get my text into the database?

Many developers face this exact problem. The theory makes sense, but writing that first working code is another story. A good tutorial closes this gap. It shows you not only what to do, but in what order and with which code. You copy the code, run it, see a result, and instantly understand how the pieces fit together.

RAG tutorials explained

This article gives you four complete walkthroughs. Each tutorial is self-contained and produces working output. You start with a minimal RAG in 30 lines, make PDFs searchable, build an interactive chat, and finally use LangChain for a structured pipeline.

All tutorials run locally. You need no cloud, no API keys, and no internet connection after setup. The language model runs via Ollama, and the vector database is Chroma. More on local RAG in the overview article.

Who these tutorials are for

These tutorials are for developers implementing RAG for the first time. You need basic Python knowledge: variables, functions, loops. Machine learning background isn’t required. If you’ve written a Python script before and run it from the command line, that’s enough.

Even if you’ve tried RAG already but aren’t sure how the pieces connect, this article will help. The tutorials are structured so you can work through each one independently or use them as templates for your own projects.

Key RAG tutorial terms

TermMeaning
TutorialStep-by-step guide with explained code
WalkthroughComplete run-through of an example from start to finish
PipelineSequence of processing steps that data passes through
StackCombination of tools and libraries used together
OllamaLocal server for running language models
ChromaLocal vector database for embeddings
LangChainFramework that connects RAG components
PythonProgramming language used in all tutorials
Virtual EnvironmentIsolated Python environment for packages
RequirementsList of required packages and versions

Tutorial 1: Minimal RAG in 30 lines

This tutorial shows the smallest possible RAG. You store three text passages in Chroma, ask a question, and get an answer from the language model. No framework, no abstraction, just core logic.

Prerequisites: Ollama is running and the llama3.2 model is downloaded. You also need the nomic-embed-text embedding model. More on embedding models in Embedding Models.

import chromadb
import requests

# Start Chroma and create collection
client = chromadb.PersistentClient(path="./minrag_db")
collection = client.get_or_create_collection("docs")

# Add documents
texts = [
    "Python is a programming language with dynamic typing.",
    "Ollama runs language models locally on your own computer.",
    "Chroma is a vector database for embeddings."
]
collection.add(
    documents=texts,
    ids=["1", "2", "3"]
)

# Ask question and find relevant passages
frage = "What is Ollama?"
results = collection.query(query_texts=[frage], n_results=2)
kontext = "\n".join(results["documents"][0])

# Build prompt and send to Ollama
prompt = f"Answer briefly in English. Use only this context:\n{kontext}\n\nQuestion: {frage}"
response = requests.post(
    "http://localhost:11434/api/generate",
    json={"model": "llama3.2", "prompt": prompt, "stream": False}
)
print(response.json()["response"])

Save the code as minrag.py and run it with python minrag.py. The output is a short answer based on the found context. Chroma uses a default embedding model internally, so you don’t need additional setup here.

The script creates a database in the minrag_db folder. On subsequent runs, Chroma loads the existing data. To add new documents, call collection.add() with new IDs.

Tutorial 2: Making PDF documents searchable

In the first tutorial, you typed in text by hand. In practice, you’ll want to search actual files, especially PDFs. This tutorial loads a PDF, splits it into passages, and makes it searchable.

You need the additional pypdf library to read PDFs. More on preparing documents in Preparing Documents.

import chromadb
import requests
from pypdf import PdfReader

# Load PDF and extract text
reader = PdfReader("example.pdf")
volltext = ""
for page in reader.pages:
    text = page.extract_text()
    if text:
        volltext += text + "\n"

# Split text into chunks (simple character-based split)
chunk_size = 500
chunk_overlap = 50
chunks = []
start = 0
while start < len(volltext):
    end = start + chunk_size
    chunks.append(volltext[start:end])
    start += chunk_size - chunk_overlap

print(f"{len(chunks)} chunks created")

# Store chunks in Chroma
client = chromadb.PersistentClient(path="./pdf_rag_db")
collection = client.get_or_create_collection("pdf_docs")
collection.add(
    documents=chunks,
    ids=[str(i) for i in range(len(chunks))]
)

# Ask question
frage = "What is this document about?"
results = collection.query(query_texts=[frage], n_results=3)
kontext = "\n\n".join(results["documents"][0])

prompt = f"Answer in English based on this context:\n{kontext}\n\nQuestion: {frage}"
response = requests.post(
    "http://localhost:11434/api/generate",
    json={"model": "llama3.2", "prompt": prompt, "stream": False}
)
print(response.json()["response"])

Replace example.pdf with the path to a real PDF file. The script reads all pages, combines the text, and splits it into 500-character passages with 50 characters of overlap. More on sensible chunking strategies in Chunking.

Make sure the PDF contains text, not just images. For scanned documents you need OCR. Details on that in the Preparing Documents article as well.

Tutorial 3: Chat with Your Documents

So far you’ve asked single questions. In practice, you’ll want to have a real conversation with context and follow-up questions. This tutorial extends the RAG setup with a chat loop that maintains conversation history.

import chromadb
import requests
import json

client = chromadb.PersistentClient(path="./chat_rag_db")
collection = client.get_or_create_collection("chat_docs")

# Load sample documents
dokumente = [
    "Der Urlaubanspruch betraegt 30 Tage im Jahr.",
    "Ueberstunden werden am Ende des Monats abgebaut oder ausgezahlt.",
    "Die Probezeit dauert sechs Monate.",
    "Homeoffice ist an zwei Tagen pro Woche moeglich."
]
collection.add(documents=dokumente, ids=["1", "2", "3", "4"])

# Conversation history
history = []

def rag_chat(frage):
    # Find relevant documents
    results = collection.query(query_texts=[frage], n_results=2)
    kontext = "\n".join(results["documents"][0])

    # Build prompt with history and context
    messages = []
    messages.append({
        "role": "system",
        "content": f"Du bist ein Assistent. Nutze diesen Kontext: {kontext}"
    })
    for eintrag in history:
        messages.append(eintrag)
    messages.append({"role": "user", "content": frage})

    # Send to Ollama (Chat API)
    response = requests.post(
        "http://localhost:11434/api/chat",
        json={"model": "llama3.2", "messages": messages, "stream": False}
    )
    antwort = response.json()["message"]["content"]

    # Update history
    history.append({"role": "user", "content": frage})
    history.append({"role": "assistant", "content": antwort})
    return antwort

# Interactive chat loop
print("Chat gestartet. Tippe 'exit' zum Beenden.")
while True:
    frage = input("Du: ")
    if frage.lower() == "exit":
        break
    antwort = rag_chat(frage)
    print(f"KI: {antwort}\n")

Save the code as chat_rag.py and run it. You can ask questions like “Wie viel Urlaub habe ich?” or “Kann ich im Homeoffice arbeiten?”. The model will search the database fresh for each query, but also consider the conversation history so far.

Ollama’s Chat API accepts a list of messages with roles. The system message provides context, while user and assistant messages form the conversation history. This lets the model understand follow-up questions like “Und wie lange ist die Probezeit?” without you having to repeat the topic.

Tutorial 4: RAG with LangChain

So far you’ve implemented each step manually. LangChain abstracts much of this away while also offering more structure and extensibility. This tutorial shows the same pipeline using LangChain.

from langchain_community.document_loaders import TextLoader
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_community.embeddings import OllamaEmbeddings
from langchain_community.vectorstores import Chroma
from langchain_community.llms import Ollama
from langchain.chains import RetrievalQA

# Load document
loader = TextLoader("notizen.txt")
docs = loader.load()

# Split into chunks
splitter = RecursiveCharacterTextSplitter(
    chunk_size=500,
    chunk_overlap=50
)
chunks = splitter.split_documents(docs)

# Embeddings and vector database
embeddings = OllamaEmbeddings(model="nomic-embed-text")
vectordb = Chroma.from_documents(
    chunks,
    embeddings,
    persist_directory="./langchain_db"
)

# Build QA chain
llm = Ollama(model="llama3.2")
qa = RetrievalQA.from_chain_type(
    llm=llm,
    chain_type="stuff",
    retriever=vectordb.as_retriever(search_kwargs={"k": 3})
)

# Ask a question
frage = "Was steht in den Notizen?"
antwort = qa.run(frage)
print(antwort)

LangChain encapsulates loading, splitting, embedding, and querying into discrete components. The RetrievalQA chain wires the retriever and LLM together automatically. You can swap components without rewriting the rest. If you want to use a different vector database instead of Chroma, change just one line.

For more complex workflows with conditionals and loops, check out LangGraph. It lets you build state machines for RAG pipelines with multiple stages, such as automatic reranking or query transformation.

Prerequisites for All Tutorials

Before running the tutorials, you need three things: Python, Ollama, and a handful of Python packages.

Install Python

Python 3.10 or newer is sufficient. Check your version with python --version. Create a virtual environment for these tutorials to keep packages from affecting your system:

python -m venv rag-env
source rag-env/bin/activate

On Windows, use rag-env\Scripts\activate instead of source.

Set Up Ollama

Download Ollama from the official website and install it. Then start the Ollama server and pull the models:

ollama pull llama3.2
ollama pull nomic-embed-text

llama3.2 is the language model, and nomic-embed-text generates embeddings. Ollama runs on port 11434 by default. For more details, see Ollama.

Install Packages

Create a requirements.txt file with these entries:

chromadb
requests
pypdf
langchain
langchain-community

Install all packages with a single command:

pip install -r requirements.txt

After that, you can run each tutorial directly.

Tips for Your Own Projects

The tutorials use example documents. For real projects, adjust a few things:

  • Documents: Replace hardcoded text with actual files. Use TextLoader, PdfReader, or DirectoryLoader for entire folders.
  • Chunk size: Test sizes between 300 and 800 characters. For technical documentation, smaller chunks are often more precise. See Chunking for more.
  • Model: llama3.2 is a solid starting point. For better answers, try mistral or qwen2.5. Larger models produce better results but need more RAM.
  • Number of results: The n_results or k parameter controls how many sections the model receives. Three to five is a good starting point.
  • Persistence: Chroma saves to disk with PersistentClient, so your data survives between runs. For testing, use EphemeralClient, which keeps everything in RAM.
  • Metadata: Add metadata to each chunk like filename, page number, or date. This lets you filter later, such as searching only in specific documents.

Common RAG Tutorial Pitfalls

  1. Ollama isn’t running: Check if Ollama is reachable with curl http://localhost:11434/api/tags. If not, start the server with ollama serve.
  2. Wrong model: If you use nomic-embed-text in LangChain, that model must also be pulled in Ollama. ollama list shows all installed models.
  3. Duplicate IDs in Chroma: Each ID can only exist once. If you run the script multiple times, use collection.upsert() instead of collection.add(), or delete the database first.
  4. PDF without text: Scanned PDFs contain only images. pypdf will extract an empty string. Use OCR libraries like pytesseract for such files.
  5. Encoding issues: Special characters can get lost with incorrect encoding. Always open files with encoding="utf-8".
  6. Too much context: If you set n_results too high, the model receives irrelevant sections. The answer becomes less accurate. Start with three results and increase only if needed.
  7. Outdated LangChain API: LangChain changes its API frequently. If an import fails, check the installed version with pip show langchain and consult the current documentation.
  8. No persistent storage: Using chromadb.Client() without PersistentClient means data disappears after the script ends. Always use a path for real projects.

Hardware, Costs, and Security in RAG Tutorials

All tutorials run on a standard laptop. The llama3.2 language model needs roughly 4 GB of RAM, while nomic-embed-text uses less than 1 GB. Chroma stores data on disk, and a few hundred documents take only a few megabytes.

Costs are zero because all tools are open source. You don’t need cloud services or paid API keys.

Security is a key advantage. Your documents, embeddings, and generated responses stay on your machine. Nothing gets sent to external servers. This matters especially with sensitive data like contracts, internal notes, or customer information. Learn more about security in local AI.

Learning to Code: RAG Tutorials in Python

Important Note

Many RAG tutorials assume Python knowledge. Once you understand the fundamentals, you can adapt examples and build your own projects more easily. IRC-Coding.de offers AI coding tutorials covering RAG, chatbots, Python, C#, and more.

Further Reading and RAG Resources

FAQ: RAG Tutorials - Common Questions

Do I need a GPU for these tutorials?

No. All tutorials run on the CPU. A GPU speeds up response generation but is not required. For llama3.2, a standard laptop processor is sufficient.

Can I run the tutorials without Ollama?

The tutorials are written for Ollama. You can replace Ollama calls with other local servers like LM Studio, but you’ll need to adjust the API calls accordingly.

Which language model is best for RAG?

For getting started, llama3.2 is a good choice. For more complex answers, try mistral or qwen2.5. If you need German-specific optimization, qwen2.5 often produces better results.

How many documents can Chroma handle?

Chroma comfortably manages tens of thousands of chunks on a regular machine. For even larger datasets or distributed setups, consider Qdrant or Weaviate instead.

Do I have to use LangChain?

No. Tutorials 1 through 3 work without LangChain. LangChain is useful when you want to swap components quickly or build complex chains. For simple applications, plain code is enough.

How do I improve answer quality?

Three levers help: better chunking, retrieving more relevant results, and using a larger model. You can also apply reranking to sort retrieved sections by relevance. See the reranking section for details.

Can I implement the tutorials in a language other than Python?

Yes. Chroma provides clients for JavaScript, Go, and Ruby. Ollama exposes an HTTP API that any language can call. The concepts are identical; only the code changes.

How do I keep the database current?

Delete old entries with collection.delete() and add new ones with collection.add(). For frequent changes, a script that detects modified files and reindexes only those is worth the effort.

Are these tutorials suitable for production use?

The tutorials are educational material. For production, you need error handling, logging, a frontend, and possibly a database like PostgreSQL instead of Chroma. The underlying concepts remain the same.

What’s the difference between RAG and fine-tuning?

RAG adds documents at runtime without modifying the model. Fine-tuning trains the model on new data. RAG is faster to set up and more flexible when documents change. Fine-tuning works better when the model needs to learn a specific style or domain terminology.

Can I search multiple file types at once?

Yes. Load PDFs with pypdf, text files with open(), Markdown with TextLoader, and web pages with WebBaseLoader. Add all chunks to the same Chroma collection. Metadata helps you identify the source later.

Sources and Further Reading

  • Chroma Documentation, chromadb.com
  • Ollama Documentation, ollama.com
  • LangChain Documentation, python.langchain.com
  • LangGraph Documentation, langchain-ai.github.io/langgraph
  • Hugging Face Embedding Models, huggingface.co
Back to Blog
Share:

Related Posts