Skip to content
BotServBotServ
LlamaIndexRAGIndexingQuery EngineAgentsPython

LlamaIndex: Framework for RAG and Agent Apps

LlamaIndex data ingestion, indexing, query engine and agents. Build RAG systems and agents with LlamaIndex.

S

schutzgeist

8 min read
LlamaIndex: Framework for RAG and Agent Apps

LlamaIndex: Framework for RAG and Agent Applications

What This Article Covers

  • What you can use LlamaIndex for and how it’s structured.
  • Core concepts: Data Ingestion, Indexing, Query Engine, and Agents.
  • How to load your own documents into an index and query them.
  • Building a local RAG workflow with Ollama and LlamaIndex.
  • Combining agents with tools and query engines.
  • Common pitfalls and mistakes when getting started.

Introduction

LlamaIndex is an open-source Python framework that connects Large Language Models with your own data. Its focus lies on data ingestion, building indexes, and answering questions from those indexes. Retrieval-Augmented Generation (RAG) plays a central role here. Instead of querying a model with only its training knowledge, LlamaIndex retrieves relevant text passages and feeds them as context to the LLM.

The architecture is modular. You load documents, split them into small chunks, compute embeddings, and store them in a vector store. You query the index through a query engine. If you want to automate decisions further, you build agents that independently select tools. This makes LlamaIndex a popular choice for data-driven applications.

Why Do You Need LlamaIndex?

Language models only know what they were trained on. Once you need to include company-specific manuals, scientific papers, or internal documentation, a standalone LLM falls short. LlamaIndex helps you prepare this information and feed it to the model strategically.

The advantage over a manual solution is the number of pre-built components. Loaders for PDFs, text files, and web pages, various splitters, multiple vector store integrations, and ready-made query engines save you significant time. Your code also stays readable because the workflow is clearly separated into loading, indexing, and querying phases.

LlamaIndex works well with local models too. If you use Ollama, your data stays on your machine. This is especially valuable for sensitive content.

LlamaIndex at a Glance

The framework consists of several layers:

  • Data Ingestion: Loading documents from various sources and converting them into a uniform format.
  • Node Parser: Splitting large documents into small chunks called nodes.
  • Embeddings: Converting text into vectors that capture semantic similarity.
  • Index: A container that manages nodes and their vectors, such as a vector store index.
  • Retriever: A component that returns the most relevant nodes for a query.
  • Query Engine: Combines retrieval and LLM invocation into an answer.
  • Agent: A ReAct or Function Agent that uses multiple tools and query engines.
  • Tool: A function or query engine that an agent can invoke.

The concept is straightforward: ingest data, find matching passages, and feed them strategically to the LLM. The result is a RAG system that delivers more current and accurate answers.

Who This Article Is For

This article is aimed at beginners with basic Python knowledge. You should understand how a language model works, but you don’t need to be a RAG expert. If you want to learn how to connect your own data to an LLM, LlamaIndex is a good starting point.

Knowledge of embeddings or vector databases helps but isn’t required. What matters more is your willingness to experiment with pipelines and APIs. If you want to dive deeper into RAG, RAG Fundamentals and Vector Databases provide additional context.

Key Terms

TermDefinition
DocumentA loaded document, such as from a PDF or text file.
NodeA small text section from a document, typically a chunk.
ChunkA contiguous text block that can be indexed separately.
EmbeddingA vector that represents the meaning of text in numbers.
Vector StoreStorage for embeddings that enables similarity searches.
IndexThe data structure that manages nodes and their vectors.
RetrieverComponent that finds relevant nodes based on a query.
Query EngineCombination of retriever, synthesizer, and LLM for answering.
AgentA system that independently selects steps and tools.
ToolA function or query engine that an agent can call.

Installation

For most examples, these packages are sufficient. Make sure you have Python 3.9 or later installed:

pip install llama-index llama-index-llms-ollama llama-index-embeddings-ollama

If you want to load PDFs, add:

pip install pypdf

Ollama should be running with at least one model like llama3.2 and an embedding model like nomic-embed-text available. You can download the models with ollama pull llama3.2 and ollama pull nomic-embed-text.

Loading and Indexing Data

The first step is loading and preparing your data. LlamaIndex provides various loaders that convert files into Document objects. Next, you split the content into manageable nodes and compute embeddings.

from llama_index.core import VectorStoreIndex, SimpleDirectoryReader, Settings
from llama_index.core.node_parser import SentenceSplitter
from llama_index.llms.ollama import Ollama
from llama_index.embeddings.ollama import OllamaEmbedding

Settings.llm = Ollama(model="llama3.2", request_timeout=120.0)
Settings.embed_model = OllamaEmbedding(model_name="nomic-embed-text")

documents = SimpleDirectoryReader("data").load_data()

parser = SentenceSplitter(chunk_size=512, chunk_overlap=50)
nodes = parser.get_nodes_from_documents(documents)

index = VectorStoreIndex(nodes)
index.storage_context.persist(persist_dir="./storage")

SimpleDirectoryReader loads all files from a directory. SentenceSplitter ensures the text is split into chunks your model can process. The VectorStoreIndex stores the associated embeddings. Using persist saves the index locally so you can reuse it later.

Query Engine and RAG

Once the index is ready, you query it. The query engine handles both retrieval and answer generation:

query_engine = index.as_query_engine()

response = query_engine.query("What does the document say about costs?")
print(response)

The query engine automatically searches for the most similar nodes and passes them as context to the LLM. You can customize this behavior, such as increasing the number of returned nodes or using only the most relevant sections:

query_engine = index.as_query_engine(similarity_top_k=5)
response = query_engine.query("What steps does the guide describe?")
print(response)

This pattern is the core of a RAG system. For more background, see Local RAG and RAG Fundamentals. To understand how embeddings work in detail, check out Embedding Models.

Agents with LlamaIndex

One of LlamaIndex’s standout features is its agent system. An agent can leverage multiple tools, such as different Query Engines, search functions, or custom Python functions. The example below demonstrates a simple ReAct agent that performs multiplication.

from llama_index.core.tools import FunctionTool
from llama_index.core.agent import ReActAgent

def multiply(a: float, b: float) -> float:
    """Multiply two numbers and return the product."""
    return a * b

tool = FunctionTool.from_defaults(fn=multiply)

agent = ReActAgent.from_tools([tool], llm=Settings.llm, verbose=True)
antwort = agent.chat("Was ist 12 mal 15?")
print(antwort)

The same agent can also receive entire Query Engines as tools. This lets you combine targeted retrieval with dynamic decision-making. For more on agent concepts, see What is an AI Agent.

LlamaIndex vs. LangChain

LangChain and LlamaIndex overlap but emphasize different areas. LlamaIndex focuses on connecting your own data to LLMs. Data Ingestion, Indexing, and Query Engine are central concepts. LangChain offers a broader toolkit for chains, tools, and agents.

For pure RAG projects, LlamaIndex often gets you up and running faster because many steps work out of the box. If you need general agent workflows or more complex chains, LlamaIndex pairs well with LangChain. Both frameworks can also be combined.

Common Pitfalls

  • Wrong model class: Ollama and ChatOllama differ. For chat messages, you need classes that support the chat format.
  • Timeout too short: Local models take longer than cloud APIs. Set request_timeout higher, around 120 seconds.
  • Chunks too large: If nodes exceed the context window, information gets truncated. Test different chunk_size values.
  • Poor retrieval results: Quality depends on your embedding model and splitter. The right model matters more than quantity.
  • Forgotten persistence: Without storage_context.persist, you lose your index when the program ends.
  • Missing error handling: API calls, especially with Ollama, can fail due to timeouts or memory issues. Plan for try-except blocks.
  • Outdated imports: LlamaIndex changes import paths frequently. Check tutorial versions and install matching packages.

Hardware, Costs, and Security

LlamaIndex itself is free and open source. Costs come from cloud APIs or your own hardware. Running locally, you pay for RAM, GPU, and electricity, but nothing per request. A 7B model often runs on 8 GB VRAM thanks to quantization. Larger models or many concurrent requests need significantly more.

RAG adds overhead through embeddings and the vector store. This is usually negligible unless your index grows enormous. Security matters when processing sensitive documents. Use local models for confidential data and monitor file system permissions. For more security advice, see Ollama.

Further Reading

FAQ

Is LlamaIndex free?

Yes, the framework is open source. Costs only arise from APIs, cloud hosting, or your own hardware.

How does LlamaIndex differ from LangChain?

LlamaIndex focuses on RAG, data ingestion, and indexing. LangChain is a more general framework for chains, tools, and agents.

Can I run LlamaIndex locally?

Yes. With Ollama and local embedding models, all data stays on your machine.

What data formats does LlamaIndex support?

PDFs, text files, Markdown, Word documents, web pages, and many more via additional loaders.

What is a Node?

A Node is a single chunk from a document that can be stored with an embedding.

How do I choose an embedding model?

Consider the language and domain of your data. nomic-embed-text works well locally and is versatile.

What is a Query Engine?

The Query Engine combines retrieval and answer generation: it finds relevant nodes and calls the LLM with the context.

How do I build an agent?

Create tools from Python functions or Query Engines and pass them to a ReActAgent.

Can I save my index?

Yes. Use storage_context.persist to save the index locally and reload it later.

Is Python 3.9 sufficient?

Yes, Python 3.9 or newer is recommended.

How fast is RAG with LlamaIndex?

It depends on your model, document size, and hardware. Small local RAG systems often respond within seconds.

Are there alternatives?

Yes. LangChain, Haystack, and DSPy are other options.

Sources

Back to Blog
Share:

Related Posts