DSPy: Programmatic Prompt Engineering
What This Article Covers
- What DSPy is and how it changes prompt engineering.
- Signatures, modules, and teleprompters as core concepts.
- Building simple predictions and chain-of-thought workflows.
- Constructing a RAG program with DSPy.
- Automatically optimizing prompts with examples.
- Common mistakes and pitfalls when working with DSPy.
Introduction
DSPy is a Python framework for programmatic prompt engineering. Instead of writing prompts by hand and iterating through countless versions, you define what a model should do and let DSPy find the best prompts and few-shot examples. The key difference from traditional approaches: you optimize programs, not individual prompt strings.
Developed by the Stanford NLP Group, the framework targets advanced users. It brings together language models, retrieval, and optimization into a unified structure. If you’ve worked with LangChain, LlamaIndex, or Haystack, you’ll recognize familiar patterns, though DSPy puts much stronger emphasis on automatic optimization.
Why Do You Need DSPy?
Prompt engineering is tedious. A small change in wording can dramatically shift the output. Once you’re juggling multiple models and tasks, things fall apart quickly. DSPy separates the desired task, called the signature, from the concrete prompt. You describe inputs and outputs, and DSPy finds the best phrasing.
This makes your code more maintainable. If you swap out a model, you don’t need to rewrite the entire prompt from scratch. DSPy adapts both the prompt and examples to fit the new model. That’s especially valuable for complex RAG systems, agents, or question-answering pipelines.
DSPy at a Glance
DSPy rests on four pillars:
- Signature: A contract describing inputs, outputs, and the task at hand.
- Module: Reusable building blocks like
Predict,ChainOfThought, orReAct. - Program: A chain of modules that work together to solve a task.
- Teleprompter: An optimizer that automatically tunes prompts and few-shot examples.
- LM: The language model DSPy calls when needed.
- RM: A retrieval model that provides context documents.
The signature is the central idea. It defines what the model receives and what it should return. DSPy generates the right prompt from the signature. You can override the prompt manually if needed, but you don’t have to.
Who Should Read This?
This article is for advanced Python developers. You should be comfortable with object-oriented programming and understand how language models work. Basic familiarity with prompt engineering and RAG helps but isn’t required.
If you’re just starting out, try LangChain or LlamaIndex first. DSPy becomes worthwhile once you want to automate recurring prompts and improve them systematically.
Key Terms
| Term | Meaning |
|---|---|
| Signature | Describes the task, inputs, and outputs for a module. |
| InputField | An input field in a signature. |
| OutputField | An output field in a signature. |
| Module | A reusable DSPy building block like Predict or ChainOfThought. |
| Predict | A module that makes a direct prediction. |
| ChainOfThought | A module that asks the model to think step by step. |
| Teleprompter | An optimizer for prompts and few-shot examples. |
| LM | The language model DSPy has configured. |
| RM | The retrieval model or retriever for context. |
| Compile | The process of optimizing a program with a teleprompter. |
Installation
Install DSPy in a virtual environment:
pip install dspy
For local use with Ollama, you don’t need an API key. Make sure Ollama is running and a model like llama3.2 is available:
ollama pull llama3.2
Install optional packages if you plan to work with ColBERT, ChromaDB, or Weaviate as retrievers.
Getting Started: Signatures and Predictions
Every DSPy program begins with a signature. It defines the task in natural language and specifies the fields.
import dspy
lm = dspy.LM(
model='ollama_chat/llama3.2',
api_base='http://localhost:11434',
api_key='',
max_tokens=512,
)
dspy.configure(lm=lm)
class GenerateAnswer(dspy.Signature):
"""Answer questions with short factoid answers."""
context = dspy.InputField(desc="may contain relevant facts")
question = dspy.InputField()
answer = dspy.OutputField(desc="often between 1 and 5 words")
qa = dspy.Predict(GenerateAnswer)
antwort = qa(
context="Das Paris befindet sich in Frankreich.",
question="In welchem Land liegt Paris?"
)
print(antwort.answer)
dspy.Predict is the simplest module type. It takes the signature and generates a call to the configured LM. The output is an object that accesses the signature’s fields.
Thinking with Chain of Thought
For trickier tasks, ChainOfThought proves invaluable. It forces the model to output intermediate reasoning before delivering the final answer.
cot = dspy.ChainOfThought(GenerateAnswer)
antwort = cot(
context="Die Hauptstadt von Frankreich ist Paris. Paris liegt am Fluss Seine.",
question="An welchem Fluss liegt die Hauptstadt von Frankreich?"
)
print(antwort.answer)
print(antwort.rationale)
The rationale field contains the reasoning the model generated internally. This boosts quality on computationally demanding or logical questions.
Building a RAG Program
DSPy shines in RAG scenarios. You combine a retrieval module with an answer generation module. This example uses a pre-built retrieval model. For local deployment, you can swap it out for your own retriever like ChromaDB or Weaviate.
import dspy
class GenerateAnswer(dspy.Signature):
"""Answer questions based on the provided context."""
context = dspy.InputField(desc="relevant passages for the question")
question = dspy.InputField()
answer = dspy.OutputField()
class RAG(dspy.Module):
def __init__(self, num_passages=3):
super().__init__()
self.num_passages = num_passages
self.generate_answer = dspy.ChainOfThought(GenerateAnswer)
def forward(self, question):
# Normally you'd use a real retriever here.
context = "Dies ist ein Beispielkontext, der in einer echten Anwendung aus einem Retriever kommt."
return self.generate_answer(context=context, question=question)
rag = RAG()
antwort = rag("Was ist DSPy?")
print(antwort.answer)
The RAG module wraps the workflow. The retriever part is simplified here. In production, you configure dspy.settings.configure(rm=retriever) and use dspy.Retrieve(k=3).
Learn more about RAG in RAG Fundamentals and Local RAG. For embeddings and vector stores, see Embedding Models and Vector Databases.
Automatically Optimize Prompts
What sets DSPy apart is the Teleprompter. You create a training set and let DSPy find the best few-shot examples and instructions.
from dspy.teleprompt import BootstrapFewShot
trainset = [
dspy.Example(
question="Wer entwickelte die Relativitätstheorie?",
answer="Albert Einstein",
).with_inputs("question"),
dspy.Example(
question="Welches ist das größte Tier?",
answer="Blauwal",
).with_inputs("question"),
]
def metrik(gold, pred, trace=None):
return gold.answer.lower() == pred.answer.lower()
teleprompter = BootstrapFewShot(metric=metrik, max_bootstrapped_demos=4)
compiled_rag = teleprompter.compile(student=RAG(), train=trainset)
antwort = compiled_rag("Was ist DSPy?")
print(antwort.answer)
BootstrapFewShot selects examples from your training set that perform best according to your metric. The result is an optimized program that uses better prompts and examples.
Other Optimizers
Beyond BootstrapFewShot, several other Teleprompters are available. MIPROv2 combines prompt optimization with example selection and works well for demanding tasks. BootstrapFewShotWithRandomSearch tries random subsets. For getting started, BootstrapFewShot is sufficient.
You don’t need a massive training dataset. DSPy often works with just a few dozen examples. What matters more is a solid metric that measures your desired behavior.
Common Pitfalls
- Outdated syntax: DSPy evolves quickly. Old tutorials use
OpenAIorColBERTv2classes that have different names in newer versions. Check your version. - Wrong model specification: For Ollama, use
ollama_chat/modeland setapi_base, notapi_base_urlorbase_url. - Missing inputs:
with_inputsis necessary so DSPy knows which fields represent the question. - Poor metric: A weak metric leads to useless optimized prompts. Invest time in your evaluation criterion.
- Too few training examples: BootstrapFewShot needs some positive examples. Optimization doesn’t work well with just one example.
- Forgotten retrieval: A RAG program without a real retriever only produces hallucinations. Configure
dspy.settings.configure(rm=...). - Incorrect signatures: Descriptions in
InputFieldandOutputFieldinfluence the prompt. Keep them precise. - Testing without compiling: Many comparisons between manual and DSPy prompts only work after the
compilestep.
Hardware, Costs, and Security
DSPy itself is free. Costs come from the language models you use. With Ollama, calls stay local. The optimization step compile can generate many model calls because DSPy tests different prompts and examples. Plan for sufficient compute time.
For larger training runs, choose a model that responds reliably and quickly. Stronger models produce better optimized prompts but need more memory. A 7B model in quantization runs on 8 GB VRAM.
Security matters once you work with sensitive training examples or retrieval sources. Keep data local and avoid public endpoints for confidential content. DSPy caches calls internally, so check caching settings if you process sensitive information.
Further Reading
- Development Tools Overview for related frameworks.
- LangChain for chains and agents.
- LlamaIndex for RAG and indexing.
- Haystack for pipelines and NLP.
- Local RAG for the full context.
- RAG Fundamentals for the theory.
- Vector Databases to choose a suitable retriever.
- Embedding Models for background on embeddings.
- Ollama for running models locally.
- What is an AI Agent for agent concepts.
FAQ
Is DSPy free?
Yes, the framework is Open Source. Costs only arise from model APIs or your own hardware.
What is programmatic prompt engineering?
It means controlling prompts and examples through code and optimization, rather than changing them by hand.
Do I need DSPy if I already use LangChain?
Not necessarily. DSPy complements other frameworks once you want to automatically optimize prompts.
What is a DSPy signature?
A signature describes which inputs a module receives and which outputs it produces.
How does Predict differ from ChainOfThought?
Predict is direct. ChainOfThought asks the model to generate reasoning before returning the answer.
What is a Teleprompter?
A Teleprompter optimizes prompts and few-shot examples based on training data and a metric.
Can I use DSPy with Ollama?
Yes. You configure dspy.LM with an Ollama endpoint.
What do I need for BootstrapFewShot?
Some examples with inputs and expected outputs, plus a function that evaluates predictions.
Is DSPy only suitable for RAG?
No. DSPy also works for classification, summarization, question-answering, and agents.
How large should the training set be?
Usually 20 to 100 examples are enough. Quality of the metric matters more than quantity.
What happens during compile?
DSPy tests different prompts and examples on your training set and generates an optimized program.
Can I connect DSPy to local retriever databases?
Yes. DSPy is retriever-agnostic. You can use ChromaDB, Weaviate, FAISS, or custom implementations.
Sources
- DSPy Documentation: https://dspy.ai/
- DSPy GitHub: https://github.com/stanfordnlp/dspy
- DSPy API: https://dspy.ai/api/
- DSPy RAG Tutorial: https://dspy.ai/tutorials/rag/
- Stanford NLP Group: https://nlp.stanford.edu/


