LlamaIndex: RAG-Focused Agent Framework
What This Article Covers
- What LlamaIndex is and how it works.
- Building agents with LlamaIndex and RAG.
- Data indices, query engines, and agent tools.
- Practical examples for document Q&A and knowledge bases.
- Best practices for RAG quality and performance.
Introduction: Understanding LlamaIndex
LlamaIndex is a framework for RAG-based agents. It connects LLMs with data indices (vector databases), query engines (search), and tools (actions). Use it when you need agents that access documents and knowledge bases.
This article targets developers building RAG agents with LlamaIndex. For foundational concepts, see Local RAG and AI Agent Frameworks.
Why Use LlamaIndex?
Imagine building an agent that answers questions about your documents. LangChain is general-purpose; LlamaIndex is specialized. It provides pre-built indices, query engines, and retrieval tools for RAG. For document-focused agents, it’s the better choice.
LlamaIndex at a Glance
LlamaIndex = RAG-focused framework. Data indices for documents, query engines for search, tools for agents. Built for document-based agents and knowledge bases.
Core idea: RAG-first design for document-focused agents.
Who Should Read This?
- RAG developers building document-based agents.
- Knowledge managers implementing document Q&A.
- Developers using LlamaIndex for agents.
- Python developers comparing RAG frameworks.
Key Concepts
- LlamaIndex - RAG framework. Use when: building RAG agents.
- RAG - Retrieval-Augmented Generation. Use when: learning the concept.
- Vector Database - Storage layer. Use when: building indices.
- Ollama - Local model server. Use when: running models locally.
- Query Engine - Search mechanism. Use when: performing retrieval.
Setup
pip install llama-index llama-index-llms-ollama llama-index-embeddings-ollama
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
from llama_index.llms.ollama import Ollama
from llama_index.embeddings.ollama import OllamaEmbedding
# Configure LLM
llm = Ollama(model="llama3.1", base_url="http://ollama:11434")
# Configure embeddings
embed_model = OllamaEmbedding(
model_name="multilingual-e5",
base_url="http://ollama:11434"
)
Example 1: Document Index
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
# Load documents
documents = SimpleDirectoryReader("./dokumente").load_data()
# Create index
index = VectorStoreIndex.from_documents(
documents,
embed_model=embed_model
)
# Query engine
query_engine = index.as_query_engine(llm=llm)
# Ask a question
response = query_engine.query("What does the contract say about termination?")
print(response)
Example 2: Agent with Tools
from llama_index.core.agent import ReActAgent
from llama_index.core.tools import FunctionTool
# Define tools
def multiply(a: int, b: int) -> int:
"""Multiply two numbers"""
return a * b
def get_weather(location: str) -> str:
"""Get weather for a location"""
return f"Weather in {location}: 22°C"
# Create tools
tools = [
FunctionTool.from_defaults(fn=multiply),
FunctionTool.from_defaults(fn=get_weather)
]
# Create agent
agent = ReActAgent.from_tools(
tools,
llm=llm,
verbose=True
)
# Use agent
response = agent.chat("What is 123 * 456?")
print(response)
Example 3: RAG Agent
from llama_index.core.agent import ReActAgent
from llama_index.core.tools import QueryEngineTool, ToolMetadata
# Query engine as tool
query_tool = QueryEngineTool(
query_engine=query_engine,
metadata=ToolMetadata(
name="document_search",
description="Search documents"
)
)
# Agent with RAG tool
agent = ReActAgent.from_tools(
[query_tool],
llm=llm,
verbose=True
)
# Agent uses RAG for answers
response = agent.chat("What does the document say about privacy?")
print(response)
Example 4: Multi-Index Agent
# Multiple indices for different document sets
index_contracts = VectorStoreIndex.from_documents(
SimpleDirectoryReader("./vertraege").load_data(),
embed_model=embed_model
)
index_invoices = VectorStoreIndex.from_documents(
SimpleDirectoryReader("./rechnungen").load_data(),
embed_model=embed_model
)
# Tools for each index
contract_tool = QueryEngineTool(
query_engine=index_contracts.as_query_engine(llm=llm),
metadata=ToolMetadata(
name="contract_search",
description="Search contracts"
)
)
invoice_tool = QueryEngineTool(
query_engine=index_invoices.as_query_engine(llm=llm),
metadata=ToolMetadata(
name="invoice_search",
description="Search invoices"
)
)
# Agent with multiple tools
agent = ReActAgent.from_tools(
[contract_tool, invoice_tool],
llm=llm
)
LlamaIndex vs. LangChain
| Aspect | LlamaIndex | LangChain |
|---|---|---|
| Focus | RAG, documents | General-purpose |
| Indices | Yes, built-in | No, manual setup |
| Query Engine | Yes, built-in | No, manual setup |
| Agents | Yes, ReAct | Yes, various |
| Complexity | Simpler for RAG | More complex, flexible |
| Best for | Document Q&A | General agents |
Security Considerations
- Data sources: Indices may contain sensitive data. Implement access controls.
- Prompt injection: Documents can contain injection attacks. See Prompt Injection.
- Audit logging: Log all agent actions. See Audit Logging.
- Permissions: Agents should only access necessary indices. See Tool Permissions.
Common Pitfalls
- Too many documents: Large indices slow down queries. Use chunking and filtering.
- Poor embeddings: Wrong embedding model means poor search. Use multilingual-e5 for German.
- Missing sources: Agents should cite sources, not just answer. Include retrieval context.
- Hallucinations: Without good retrieval, models hallucinate. Test RAG quality.
- Too many tools: More than 5-7 tools overwhelm the model.
Further Reading
- LlamaIndex - GitHub.
- Local RAG - RAG fundamentals.
- AI Agent Frameworks - Overview.
- LangChain - Alternative.
- Semantic Kernel - Alternative.
- Ollama - Model server.
Key Takeaways:
- LlamaIndex = RAG-focused framework for document-based agents.
- Data indices, query engines, and tools for agents.
- For document Q&A, knowledge bases, and RAG applications.
- Simpler for RAG than LangChain, more specialized.
- For general agents: use LangChain or CrewAI.
FAQ
What is LlamaIndex?
LlamaIndex or LangChain?
What are data indices?
What is a Query Engine?
Which embedding model?
Are my documents safe?
What does LlamaIndex cost?
How performant is LlamaIndex?
Sources and Further Reading
- LlamaIndex - GitHub.
- LlamaIndex Docs - Documentation.
- RAG Guide - RAG techniques.


