LangChain: Framework for LLM Applications
What this article covers
- What LangChain is and how you can use it.
- Core concepts: chains, prompt templates, retrievers, agents, and more.
- Why a framework often beats working directly with APIs.
- Installation and your first working example with Ollama.
- How RAG and agents work with LangChain.
- Common pitfalls and mistakes when getting started.
Introduction: Understanding LangChain
LangChain is an open-source framework for Python and JavaScript that lets you build applications powered by Large Language Models. It provides pre-built components you can chain together. A chain is an executable workflow consisting of a prompt, a model call, and response processing.
Many repetitive tasks are already solved: reading PDFs, searching a vector store, calling a search engine, or converting model output into clean structured data. You write less boilerplate code and focus on your own logic.
LangGraph, which builds on top of LangChain, gets its own article at LangGraph so we don’t cover the same ground twice.
Why do you need LangChain?
Imagine you want to connect a local LLM with a PDF and web search. Without a framework, you’d stitch together multiple libraries with different APIs. LangChain handles much of this glue code for you. You load a PDF with a Document Loader, split it into chunks, store those in a Vector Store, and integrate the results into a chain. For web search, you pick an appropriate tool from the package.
LangChain is also portable. Many components work with both cloud APIs and local models. If you start with Ollama and later want to try a different model, you often change just one line.
LangChain at a glance
At its core, LangChain offers these components:
- Models: Abstractions for chat models, LLMs, and embeddings.
- Prompt Templates: Templates that format inputs consistently for the model.
- Document Loaders: Modules that convert file formats into text.
- Text Splitters: Tools to break long texts into chunks.
- Vector Stores: Databases for semantic search.
- Retrievers: Components that fetch relevant documents from a vector store.
- Output Parsers: Tools that structure model responses into fixed formats.
- Chains: Linked steps executed like a pipeline.
- Agents: Components that decide independently which tools to use.
- Memory: Ways to store context across multiple messages.
The newest style is LCEL, LangChain Expression Language. You chain components with the pipe operator |, much like Unix pipes.
Who is LangChain for?
LangChain targets developers who want to integrate language models into applications. You should be comfortable with Python or JavaScript and not shy away from pip or npm. Prior knowledge of RAG or agents makes the learning curve easier but isn’t required. If you prefer working visually, LangFlow is a better fit. For complex state graphs, LangGraph is worth exploring.
Key concepts in LangChain
| Term | Meaning |
|---|---|
| Chain | A connected sequence of steps, typically prompt, model, and parser. |
| Prompt Template | A template that formats inputs consistently for the model. |
| Retriever | A component that returns matching documents from a vector store. |
| Document Loader | A module that converts PDFs, text files, or web pages into documents. |
| Output Parser | Converts model output into a desired structure like lists or objects. |
| Agent | A mechanism that decides which tools and steps are needed. |
| Tool | A function that an agent or chain can call, such as a web search. |
| Memory | Storage for past messages or intermediate results. |
| Callback | A hook to log events during execution. |
| LCEL | LangChain Expression Language, pipe notation for chaining components. |
LangChain vs. direct API calls
For simple prompt-and-response interactions, calling the API directly is often faster. Once you combine multiple steps, you hit these friction points:
- Repetition: You implement prompt formatting, retry logic, and error handling yourself.
- Format differences: Each model provider has a different message format. LangChain abstracts these away.
- Data integration: Connecting PDFs, web pages, and vector stores creates a lot of overhead.
- Maintainability: A chain of pre-built components is easier to read and modify.
LangChain isn’t mandatory, but it’s a practical tool once your application grows beyond simple prompts.
Installation
For most examples, these commands suffice:
pip install langchain langchain-community langchain-ollama
For PDF support, add:
pip install pypdf
For Chroma as your vector store:
pip install chromadb
For JavaScript:
npm install langchain
Use Python 3.9 or newer. Ollama must be installed with at least one model available. See Ollama for details.
Your first example
A simple starting point is a chain with a prompt template, model, and output parser:
from langchain_ollama import ChatOllama
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
model = ChatOllama(model="llama3.2")
prompt = ChatPromptTemplate.from_messages([
("system", "Du bist ein freundlicher Assistent, der kurz und klar antwortet."),
("human", "Erkläre mir {thema} in drei Sätzen."),
])
parser = StrOutputParser()
chain = prompt | model | parser
antwort = chain.invoke({"thema": "LangChain"})
print(antwort)
The prompt template accepts the thema variable, the model responds, and the parser returns plain text.
RAG with LangChain
RAG, Retrieval-Augmented Generation, is a common use case. The workflow is:
- Load documents with a Document Loader.
- Split them into small chunks.
- Compute embeddings and store them in a Vector Store.
- When answering a query, retrieve the most relevant chunks.
- Combine the retrieved chunks and user question in a chain.
A minimal example with Ollama, a PDF, and Chroma:
from langchain_community.document_loaders import PyPDFLoader
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_ollama import OllamaEmbeddings
from langchain_community.vectorstores import Chroma
from langchain_ollama import ChatOllama
from langchain.chains.combine_documents import create_stuff_documents_chain
from langchain.chains import create_retrieval_chain
from langchain_core.prompts import ChatPromptTemplate
loader = PyPDFLoader("dokument.pdf")
docs = loader.load()
splitter = RecursiveCharacterTextSplitter(chunk_size=500, chunk_overlap=50)
chunks = splitter.split_documents(docs)
embeddings = OllamaEmbeddings(model="nomic-embed-text")
vectorstore = Chroma.from_documents(chunks, embeddings)
retriever = vectorstore.as_retriever()
model = ChatOllama(model="llama3.2")
prompt = ChatPromptTemplate.from_messages([
("system", "Beantworte die Frage ausschließlich anhand des bereitgestellten Kontexts."),
("human", "Kontext: {context}\n\nFrage: {input}"),
])
combine_chain = create_stuff_documents_chain(model, prompt)
rag_chain = create_retrieval_chain(retriever, combine_chain)
antwort = rag_chain.invoke({"input": "Was steht im Dokument zum Thema Kosten?"})
print(antwort["answer"])
You can adapt this pattern to different file types, embedding models, and vector stores. For more background, see Local RAG and RAG Fundamentals.
Agents with LangChain
Agents decide for themselves which tools to invoke. Here’s a simple ReAct example:
from langchain_ollama import ChatOllama
from langchain_community.tools import DuckDuckGoSearchRun
from langchain.agents import create_react_agent, AgentExecutor
from langchain import hub
model = ChatOllama(model="llama3.2")
tools = [DuckDuckGoSearchRun()]
prompt = hub.pull("hwchase17/react")
agent = create_react_agent(model, tools, prompt)
agent_executor = AgentExecutor(agent=agent, tools=tools, verbose=True)
response = agent_executor.invoke({"input": "What's the weather like in Berlin today?"})
print(response["output"])
Not every model handles tool calling reliably. Qwen, Mistral, and Command R tend to work better than very small variants. Learn more in Tool-Calling and What is an AI Agent.
LangChain and Ollama
LangChain interacts with Ollama through three main classes: ChatOllama, Ollama, and OllamaEmbeddings.
ChatOllama: For chat prompts with roles likesystem,human, andai.Ollama: For text completion models.OllamaEmbeddings: For local embeddings in RAG projects.
Here’s an example using ChatOllama:
from langchain_ollama import ChatOllama
model = ChatOllama(
model="llama3.2",
temperature=0.7,
base_url="http://localhost:11434"
)
response = model.invoke("Explain LangChain in one sentence.")
print(response.content)
Make sure Ollama is running and your chosen model is already downloaded. For more on the API, see Ollama API.
LangChain vs. LangGraph vs. LangFlow
| Aspect | LangChain | LangGraph | LangFlow |
|---|---|---|---|
| Level | Framework for chains and integrations. | Graph engine built on LangChain. | Visual editor for LangChain workflows. |
| Code | Python or JavaScript. | Python or JavaScript. | No-code to low-code. |
| Strengths | Fast prototyping, RAG, tool calling. | State-based workflows, loops, multi-agent. | Clear visual pipelines. |
| Control Flow | Linear or conditional via chains. | As a graph with nodes and edges. | Connected building blocks. |
| Learning Curve | Moderate. | Somewhat steeper. | Low if you prefer not to write code. |
| Ideal For | First projects and established pipelines. | Complex decision trees. | Experimentation and demos. |
We cover LangGraph separately since it introduces its own concepts like state, nodes, and conditional edges. LangFlow works well for understanding workflows without deep code knowledge. LangChain remains the common foundation for both.
Common LangChain Pitfalls
- Using the wrong model class:
OllamaandChatOllamaare not interchangeable. For chat prompts, useChatOllama. - Ignoring token limits: Long chains with extensive context can exceed the context window.
- Choosing the wrong output parser:
JsonOutputParserfails if the model doesn’t return valid JSON.StrOutputParseris safer for getting started. - Poor retriever results: RAG quality depends heavily on chunking strategy, embedding model choice, and top-k parameters.
- Using weak models for agents: Small models often fail at reliable tool calling.
- Version confusion: Older tutorials reference outdated import paths like
langchain.chat_models. - Missing error handling: API calls can fail. Use try-except blocks or retry logic from the start.
- Too many dependencies: Install only the integrations you actually need.
Hardware, Costs, and Security with LangChain
LangChain itself is free and open source. Costs come from models and infrastructure. Running locally, you pay for hardware and electricity, not per call. Cloud APIs charge per token but run on shared hardware.
For local deployment: larger models need more RAM or VRAM. A 7B model often runs on 8 GB VRAM thanks to quantization, but bigger variants demand significantly more. RAG adds overhead through embeddings and vector storage, though it remains relatively affordable.
Security matters once agents can perform web searches or read files. Check permissions and sandboxing before deploying anything publicly.
Further Reading and Resources on LangChain
- Development Tools Overview for related frameworks.
- LangGraph for complex state-based workflows.
- LangFlow for the visual approach.
- What is an AI Agent for agent concepts.
- Tool-Calling for external functions.
- Local RAG and RAG Fundamentals for data-driven applications.
- Ollama and Ollama API for local model deployment.
FAQ: LangChain - Common Questions
Is LangChain free?
Yes. The framework is open source. Costs arise only from models, hosting, or additional services.
Do I need Python knowledge?
Yes for most tutorials. There’s also a JavaScript version, but Python is more widely supported.
Can I run LangChain entirely locally?
Yes. With Ollama, local models, and local vector stores like Chroma or FAISS, it’s fully possible.
What’s the difference between LangChain and LlamaIndex?
LangChain covers chains, tools, and agents. LlamaIndex focuses more on RAG and working with your own data.
Which model works best with LangChain?
For simple chains, small models like Llama 3.2 suffice. For RAG and agents, reach for stronger variants like Mistral, Qwen, or Command R.
How do I learn best?
Build your first chain example, extend it with RAG, then try agents. BotServ.de offers suitable introductory articles.
Why are there so many langchain-* packages?
LangChain was split to keep dependencies clean. langchain is the core, langchain-community contains integrations, and langchain-ollama adds Ollama support.
What is LCEL?
LCEL stands for LangChain Expression Language. It lets you chain components together using the pipe operator.
Can I use LangChain with my own fine-tuned model?
Yes, as long as your model is reachable via a compatible interface.
Are there alternatives to LangChain?
Yes. Alongside LlamaIndex, LangGraph, and LangFlow, there’s CrewAI, Haystack, and custom implementations.
How stable is the API?
LangChain evolves quickly. Make sure tutorials match your installed version to avoid outdated imports.
Sources and Further Reading
- LangChain Documentation: https://python.langchain.com/
- LangChain JS/TS Documentation: https://js.langchain.com/
- LangChain GitHub Repository: https://github.com/langchain-ai/langchain
- Ollama LangChain Integration: https://python.langchain.com/docs/integrations/chat/ollama/
- LangGraph: https://langchain-ai.github.io/langgraph/
- LangFlow: https://www.langflow.org/


