Skip to content
BotServBotServ
LlamaIndexRAGAgent FrameworkData IndexesQuery Engine

LlamaIndex: RAG-Focused Agent Framework

LlamaIndex for AI agents. RAG framework, data indexes, query engines, and practical examples.

S

schutzgeist

4 min read
LlamaIndex: RAG-Focused Agent Framework

LlamaIndex: RAG-Focused Agent Framework

What This Article Covers

  • What LlamaIndex is and how it works.
  • Building agents with LlamaIndex and RAG.
  • Data indices, query engines, and agent tools.
  • Practical examples for document Q&A and knowledge bases.
  • Best practices for RAG quality and performance.

Introduction: Understanding LlamaIndex

LlamaIndex is a framework for RAG-based agents. It connects LLMs with data indices (vector databases), query engines (search), and tools (actions). Use it when you need agents that access documents and knowledge bases.

This article targets developers building RAG agents with LlamaIndex. For foundational concepts, see Local RAG and AI Agent Frameworks.

Why Use LlamaIndex?

Imagine building an agent that answers questions about your documents. LangChain is general-purpose; LlamaIndex is specialized. It provides pre-built indices, query engines, and retrieval tools for RAG. For document-focused agents, it’s the better choice.

LlamaIndex at a Glance

LlamaIndex = RAG-focused framework. Data indices for documents, query engines for search, tools for agents. Built for document-based agents and knowledge bases.

Core idea: RAG-first design for document-focused agents.

Who Should Read This?

  • RAG developers building document-based agents.
  • Knowledge managers implementing document Q&A.
  • Developers using LlamaIndex for agents.
  • Python developers comparing RAG frameworks.

Key Concepts

  • LlamaIndex - RAG framework. Use when: building RAG agents.
  • RAG - Retrieval-Augmented Generation. Use when: learning the concept.
  • Vector Database - Storage layer. Use when: building indices.
  • Ollama - Local model server. Use when: running models locally.
  • Query Engine - Search mechanism. Use when: performing retrieval.

Setup

pip install llama-index llama-index-llms-ollama llama-index-embeddings-ollama
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
from llama_index.llms.ollama import Ollama
from llama_index.embeddings.ollama import OllamaEmbedding

# Configure LLM
llm = Ollama(model="llama3.1", base_url="http://ollama:11434")

# Configure embeddings
embed_model = OllamaEmbedding(
    model_name="multilingual-e5",
    base_url="http://ollama:11434"
)

Example 1: Document Index

from llama_index.core import VectorStoreIndex, SimpleDirectoryReader

# Load documents
documents = SimpleDirectoryReader("./dokumente").load_data()

# Create index
index = VectorStoreIndex.from_documents(
    documents,
    embed_model=embed_model
)

# Query engine
query_engine = index.as_query_engine(llm=llm)

# Ask a question
response = query_engine.query("What does the contract say about termination?")
print(response)

Example 2: Agent with Tools

from llama_index.core.agent import ReActAgent
from llama_index.core.tools import FunctionTool

# Define tools
def multiply(a: int, b: int) -> int:
    """Multiply two numbers"""
    return a * b

def get_weather(location: str) -> str:
    """Get weather for a location"""
    return f"Weather in {location}: 22°C"

# Create tools
tools = [
    FunctionTool.from_defaults(fn=multiply),
    FunctionTool.from_defaults(fn=get_weather)
]

# Create agent
agent = ReActAgent.from_tools(
    tools,
    llm=llm,
    verbose=True
)

# Use agent
response = agent.chat("What is 123 * 456?")
print(response)

Example 3: RAG Agent

from llama_index.core.agent import ReActAgent
from llama_index.core.tools import QueryEngineTool, ToolMetadata

# Query engine as tool
query_tool = QueryEngineTool(
    query_engine=query_engine,
    metadata=ToolMetadata(
        name="document_search",
        description="Search documents"
    )
)

# Agent with RAG tool
agent = ReActAgent.from_tools(
    [query_tool],
    llm=llm,
    verbose=True
)

# Agent uses RAG for answers
response = agent.chat("What does the document say about privacy?")
print(response)

Example 4: Multi-Index Agent

# Multiple indices for different document sets
index_contracts = VectorStoreIndex.from_documents(
    SimpleDirectoryReader("./vertraege").load_data(),
    embed_model=embed_model
)

index_invoices = VectorStoreIndex.from_documents(
    SimpleDirectoryReader("./rechnungen").load_data(),
    embed_model=embed_model
)

# Tools for each index
contract_tool = QueryEngineTool(
    query_engine=index_contracts.as_query_engine(llm=llm),
    metadata=ToolMetadata(
        name="contract_search",
        description="Search contracts"
    )
)

invoice_tool = QueryEngineTool(
    query_engine=index_invoices.as_query_engine(llm=llm),
    metadata=ToolMetadata(
        name="invoice_search",
        description="Search invoices"
    )
)

# Agent with multiple tools
agent = ReActAgent.from_tools(
    [contract_tool, invoice_tool],
    llm=llm
)

LlamaIndex vs. LangChain

AspectLlamaIndexLangChain
FocusRAG, documentsGeneral-purpose
IndicesYes, built-inNo, manual setup
Query EngineYes, built-inNo, manual setup
AgentsYes, ReActYes, various
ComplexitySimpler for RAGMore complex, flexible
Best forDocument Q&AGeneral agents

Security Considerations

  • Data sources: Indices may contain sensitive data. Implement access controls.
  • Prompt injection: Documents can contain injection attacks. See Prompt Injection.
  • Audit logging: Log all agent actions. See Audit Logging.
  • Permissions: Agents should only access necessary indices. See Tool Permissions.

Common Pitfalls

  • Too many documents: Large indices slow down queries. Use chunking and filtering.
  • Poor embeddings: Wrong embedding model means poor search. Use multilingual-e5 for German.
  • Missing sources: Agents should cite sources, not just answer. Include retrieval context.
  • Hallucinations: Without good retrieval, models hallucinate. Test RAG quality.
  • Too many tools: More than 5-7 tools overwhelm the model.

Further Reading

Key Takeaways:

  • LlamaIndex = RAG-focused framework for document-based agents.
  • Data indices, query engines, and tools for agents.
  • For document Q&A, knowledge bases, and RAG applications.
  • Simpler for RAG than LangChain, more specialized.
  • For general agents: use LangChain or CrewAI.

FAQ

What is LlamaIndex?

A RAG-focused framework for AI agents: data indices for documents, query engines for search, tools for agents. Built for document-based agents.

LlamaIndex or LangChain?

LlamaIndex for RAG and document Q&A. LangChain for general agents and complex workflows. Both can work together.

What are data indices?

Structures for document search: VectorStoreIndex for semantic search, KeywordTableIndex for keywords, and others. They enable RAG.

What is a Query Engine?

The search mechanism for an index. It finds relevant documents, generates context, and returns answers with citations.

Which embedding model?

multilingual-e5 or bge-m3 for German documents. nomic-embed-text for English. For long documents: bge-m3 or gte-Qwen2.

Are my documents safe?

Yes, if you use local models and local indices. All data stays on your server. Cloud APIs send data externally.

What does LlamaIndex cost?

Free. LlamaIndex is open source. Local models (Ollama): hardware costs only. Cloud APIs: API costs apply.

How performant is LlamaIndex?

Good for medium document volumes (thousands of documents). For larger volumes: chunking strategy, filtering, and quality embeddings matter.

Sources and Further Reading

Back to Blog
Share:

Related Posts