Skip to content
BotServBotServ
DSPyPrompt EngineeringOptimizationLLMPythonRAG

DSPy: Programmatic Prompt Engineering

DSPy optimizes prompts and chains automatically. Programmatic LLM engineering with modules and teleprompters.

S

schutzgeist

7 min read
DSPy: Programmatic Prompt Engineering

DSPy: Programmatic Prompt Engineering

What This Article Covers

  • What DSPy is and how it changes prompt engineering.
  • Signatures, modules, and teleprompters as core concepts.
  • Building simple predictions and chain-of-thought workflows.
  • Constructing a RAG program with DSPy.
  • Automatically optimizing prompts with examples.
  • Common mistakes and pitfalls when working with DSPy.

Introduction

DSPy is a Python framework for programmatic prompt engineering. Instead of writing prompts by hand and iterating through countless versions, you define what a model should do and let DSPy find the best prompts and few-shot examples. The key difference from traditional approaches: you optimize programs, not individual prompt strings.

Developed by the Stanford NLP Group, the framework targets advanced users. It brings together language models, retrieval, and optimization into a unified structure. If you’ve worked with LangChain, LlamaIndex, or Haystack, you’ll recognize familiar patterns, though DSPy puts much stronger emphasis on automatic optimization.

Why Do You Need DSPy?

Prompt engineering is tedious. A small change in wording can dramatically shift the output. Once you’re juggling multiple models and tasks, things fall apart quickly. DSPy separates the desired task, called the signature, from the concrete prompt. You describe inputs and outputs, and DSPy finds the best phrasing.

This makes your code more maintainable. If you swap out a model, you don’t need to rewrite the entire prompt from scratch. DSPy adapts both the prompt and examples to fit the new model. That’s especially valuable for complex RAG systems, agents, or question-answering pipelines.

DSPy at a Glance

DSPy rests on four pillars:

  • Signature: A contract describing inputs, outputs, and the task at hand.
  • Module: Reusable building blocks like Predict, ChainOfThought, or ReAct.
  • Program: A chain of modules that work together to solve a task.
  • Teleprompter: An optimizer that automatically tunes prompts and few-shot examples.
  • LM: The language model DSPy calls when needed.
  • RM: A retrieval model that provides context documents.

The signature is the central idea. It defines what the model receives and what it should return. DSPy generates the right prompt from the signature. You can override the prompt manually if needed, but you don’t have to.

Who Should Read This?

This article is for advanced Python developers. You should be comfortable with object-oriented programming and understand how language models work. Basic familiarity with prompt engineering and RAG helps but isn’t required.

If you’re just starting out, try LangChain or LlamaIndex first. DSPy becomes worthwhile once you want to automate recurring prompts and improve them systematically.

Key Terms

TermMeaning
SignatureDescribes the task, inputs, and outputs for a module.
InputFieldAn input field in a signature.
OutputFieldAn output field in a signature.
ModuleA reusable DSPy building block like Predict or ChainOfThought.
PredictA module that makes a direct prediction.
ChainOfThoughtA module that asks the model to think step by step.
TeleprompterAn optimizer for prompts and few-shot examples.
LMThe language model DSPy has configured.
RMThe retrieval model or retriever for context.
CompileThe process of optimizing a program with a teleprompter.

Installation

Install DSPy in a virtual environment:

pip install dspy

For local use with Ollama, you don’t need an API key. Make sure Ollama is running and a model like llama3.2 is available:

ollama pull llama3.2

Install optional packages if you plan to work with ColBERT, ChromaDB, or Weaviate as retrievers.

Getting Started: Signatures and Predictions

Every DSPy program begins with a signature. It defines the task in natural language and specifies the fields.

import dspy

lm = dspy.LM(
    model='ollama_chat/llama3.2',
    api_base='http://localhost:11434',
    api_key='',
    max_tokens=512,
)
dspy.configure(lm=lm)

class GenerateAnswer(dspy.Signature):
    """Answer questions with short factoid answers."""

    context = dspy.InputField(desc="may contain relevant facts")
    question = dspy.InputField()
    answer = dspy.OutputField(desc="often between 1 and 5 words")

qa = dspy.Predict(GenerateAnswer)

antwort = qa(
    context="Das Paris befindet sich in Frankreich.",
    question="In welchem Land liegt Paris?"
)
print(antwort.answer)

dspy.Predict is the simplest module type. It takes the signature and generates a call to the configured LM. The output is an object that accesses the signature’s fields.

Thinking with Chain of Thought

For trickier tasks, ChainOfThought proves invaluable. It forces the model to output intermediate reasoning before delivering the final answer.

cot = dspy.ChainOfThought(GenerateAnswer)

antwort = cot(
    context="Die Hauptstadt von Frankreich ist Paris. Paris liegt am Fluss Seine.",
    question="An welchem Fluss liegt die Hauptstadt von Frankreich?"
)
print(antwort.answer)
print(antwort.rationale)

The rationale field contains the reasoning the model generated internally. This boosts quality on computationally demanding or logical questions.

Building a RAG Program

DSPy shines in RAG scenarios. You combine a retrieval module with an answer generation module. This example uses a pre-built retrieval model. For local deployment, you can swap it out for your own retriever like ChromaDB or Weaviate.

import dspy

class GenerateAnswer(dspy.Signature):
    """Answer questions based on the provided context."""

    context = dspy.InputField(desc="relevant passages for the question")
    question = dspy.InputField()
    answer = dspy.OutputField()

class RAG(dspy.Module):
    def __init__(self, num_passages=3):
        super().__init__()
        self.num_passages = num_passages
        self.generate_answer = dspy.ChainOfThought(GenerateAnswer)

    def forward(self, question):
        # Normally you'd use a real retriever here.
        context = "Dies ist ein Beispielkontext, der in einer echten Anwendung aus einem Retriever kommt."
        return self.generate_answer(context=context, question=question)

rag = RAG()
antwort = rag("Was ist DSPy?")
print(antwort.answer)

The RAG module wraps the workflow. The retriever part is simplified here. In production, you configure dspy.settings.configure(rm=retriever) and use dspy.Retrieve(k=3).

Learn more about RAG in RAG Fundamentals and Local RAG. For embeddings and vector stores, see Embedding Models and Vector Databases.

Automatically Optimize Prompts

What sets DSPy apart is the Teleprompter. You create a training set and let DSPy find the best few-shot examples and instructions.

from dspy.teleprompt import BootstrapFewShot

trainset = [
    dspy.Example(
        question="Wer entwickelte die Relativitätstheorie?",
        answer="Albert Einstein",
    ).with_inputs("question"),
    dspy.Example(
        question="Welches ist das größte Tier?",
        answer="Blauwal",
    ).with_inputs("question"),
]

def metrik(gold, pred, trace=None):
    return gold.answer.lower() == pred.answer.lower()

teleprompter = BootstrapFewShot(metric=metrik, max_bootstrapped_demos=4)
compiled_rag = teleprompter.compile(student=RAG(), train=trainset)

antwort = compiled_rag("Was ist DSPy?")
print(antwort.answer)

BootstrapFewShot selects examples from your training set that perform best according to your metric. The result is an optimized program that uses better prompts and examples.

Other Optimizers

Beyond BootstrapFewShot, several other Teleprompters are available. MIPROv2 combines prompt optimization with example selection and works well for demanding tasks. BootstrapFewShotWithRandomSearch tries random subsets. For getting started, BootstrapFewShot is sufficient.

You don’t need a massive training dataset. DSPy often works with just a few dozen examples. What matters more is a solid metric that measures your desired behavior.

Common Pitfalls

  • Outdated syntax: DSPy evolves quickly. Old tutorials use OpenAI or ColBERTv2 classes that have different names in newer versions. Check your version.
  • Wrong model specification: For Ollama, use ollama_chat/model and set api_base, not api_base_url or base_url.
  • Missing inputs: with_inputs is necessary so DSPy knows which fields represent the question.
  • Poor metric: A weak metric leads to useless optimized prompts. Invest time in your evaluation criterion.
  • Too few training examples: BootstrapFewShot needs some positive examples. Optimization doesn’t work well with just one example.
  • Forgotten retrieval: A RAG program without a real retriever only produces hallucinations. Configure dspy.settings.configure(rm=...).
  • Incorrect signatures: Descriptions in InputField and OutputField influence the prompt. Keep them precise.
  • Testing without compiling: Many comparisons between manual and DSPy prompts only work after the compile step.

Hardware, Costs, and Security

DSPy itself is free. Costs come from the language models you use. With Ollama, calls stay local. The optimization step compile can generate many model calls because DSPy tests different prompts and examples. Plan for sufficient compute time.

For larger training runs, choose a model that responds reliably and quickly. Stronger models produce better optimized prompts but need more memory. A 7B model in quantization runs on 8 GB VRAM.

Security matters once you work with sensitive training examples or retrieval sources. Keep data local and avoid public endpoints for confidential content. DSPy caches calls internally, so check caching settings if you process sensitive information.

Further Reading

FAQ

Is DSPy free?

Yes, the framework is Open Source. Costs only arise from model APIs or your own hardware.

What is programmatic prompt engineering?

It means controlling prompts and examples through code and optimization, rather than changing them by hand.

Do I need DSPy if I already use LangChain?

Not necessarily. DSPy complements other frameworks once you want to automatically optimize prompts.

What is a DSPy signature?

A signature describes which inputs a module receives and which outputs it produces.

How does Predict differ from ChainOfThought?

Predict is direct. ChainOfThought asks the model to generate reasoning before returning the answer.

What is a Teleprompter?

A Teleprompter optimizes prompts and few-shot examples based on training data and a metric.

Can I use DSPy with Ollama?

Yes. You configure dspy.LM with an Ollama endpoint.

What do I need for BootstrapFewShot?

Some examples with inputs and expected outputs, plus a function that evaluates predictions.

Is DSPy only suitable for RAG?

No. DSPy also works for classification, summarization, question-answering, and agents.

How large should the training set be?

Usually 20 to 100 examples are enough. Quality of the metric matters more than quantity.

What happens during compile?

DSPy tests different prompts and examples on your training set and generates an optimized program.

Can I connect DSPy to local retriever databases?

Yes. DSPy is retriever-agnostic. You can use ChromaDB, Weaviate, FAISS, or custom implementations.

Sources

Back to Blog
Share:

Related Posts