aiagent.org logo aiagent.org
⚡ Declarative AI Framework

DSPy: Algorithmic Prompt Compilation

Replace brittle, manual prompt engineering with algorithmic prompt optimization. Learn DSPy Signatures, ChainOfThought modules, ReAct agents, and automatic teleprompters (BootstrapFewShot & MIPROv2) that optimize your system prompts against quantitative metrics.

Core Paradigm: Programming, Not Prompting
Optimizers: BootstrapFewShot & MIPROv2
Impact: +20–35% Benchmark Accuracy Boost

01 Why Manual Prompt Engineering Fails at Scale

Philosophy

When you handcraft long prompt strings with specific phrasing, every model update (e.g. from GPT-5 to GPT-6 or Claude 3.5 to Claude Opus 5.5) breaks your formatting. DSPy (created by Stanford NLP) separates the logic of your program from the textual prompt instructions:

Traditional Prompt Engineering
  • Fiddling with adjectives ("You are an elite, world-class expert...").
  • Hardcoding few-shot examples that go stale.
  • Breaks completely when changing foundation models or temperature.
  • No systematic way to hill-climb toward optimal accuracy.
The DSPy Declarative Paradigm
  • Declare inputs and outputs using typed Signatures.
  • Build reusable pipelines using modular Modules.
  • Use an Optimizer (Teleprompter) to synthesize few-shot examples and system instructions based on your validation set.

02 Building a DSPy Pipeline with Self-Optimization

Python 3.10+

Here is a complete DSPy pipeline that defines an Agent Signature, evaluates answers with a custom metric, and compiles optimal prompts automatically using BootstrapFewShot:

# Install: pip install dspy-ai
import dspy
from dspy.teleprompt import BootstrapFewShot

# 1. Configure default LLM provider
lm = dspy.LM('openai/gpt-6-sol', api_key=os.getenv('OPENAI_API_KEY'))
dspy.configure(lm=lm)

# 2. Define a typed Signature (declarative input/output contract)
class FinancialAnalysisSignature(dspy.Signature):
    """Analyze enterprise earnings reports and detect liquidity risks."""
    quarterly_filing = dspy.InputField(desc="Raw excerpt from 10-Q filing")
    question = dspy.InputField(desc="Financial query to investigate")
    reasoning = dspy.OutputField(desc="Step-by-step financial ratio calculation")
    risk_rating = dspy.OutputField(desc="Liquidity risk rating: LOW, MEDIUM, or HIGH")

# 3. Define the Agent Module
class FinancialRiskAgent(dspy.Module):
    def __init__(self):
        super().__init__()
        # ChainOfThought automatically adds an internal reasoning step
        self.prog = dspy.ChainOfThought(FinancialAnalysisSignature)

    def forward(self, quarterly_filing, question):
        return self.prog(quarterly_filing=quarterly_filing, question=question)

# 4. Define quantitative evaluation metric
def metric(gold, pred, trace=None):
    # Matches risk rating and ensures reasoning contains financial citations
    return gold.risk_rating.upper() == pred.risk_rating.upper() and len(pred.reasoning) > 50

# 5. Compile and Optimize! (Teleprompter creates optimal few-shot exemplars)
trainset = [
    dspy.Example(
        quarterly_filing="Cash reserves dropped 45% to $12M while accounts payable surged to $34M.",
        question="What is the immediate liquidity risk?",
        risk_rating="HIGH"
    ).with_inputs('quarterly_filing', 'question'),
]

teleprompter = BootstrapFewShot(metric=metric, max_bootstrapped_demos=3)
compiled_agent = teleprompter.compile(FinancialRiskAgent(), trainset=trainset)

# Execute the self-optimized agent
result = compiled_agent(
    quarterly_filing="Operating cash flow rose 18% with $400M in cash and negligible short-term debt.",
    question="Evaluate debt default risk."
)
print("Reasoning:", result.reasoning)
print("Risk Rating:", result.risk_rating)

03 MIPROv2: Multi-Prompt Instruction Optimization

Advanced

While BootstrapFewShot synthesizes few-shot examples, MIPROv2 (Multiprompt Instruction Proposal and Optimization) goes a step further: it uses an LLM to propose candidate system prompts, samples combinations with Bayesian optimization, and finds the highest-scoring instructions across your evaluation test set.

Frequently Asked Questions

How does DSPy compare to LangChain or CrewAI? →

LangChain and CrewAI are orchestration frameworks for connecting tools, chains, and multi-agent roles. DSPy is an optimization compiler. You can use DSPy to automatically discover the best prompts and exemplars for agents, and then deploy those compiled artifacts into any backend system.

How much training data do I need to compile a DSPy module? →

Surprisingly little: even 10 to 30 well-curated examples can produce double-digit accuracy gains because DSPy uses the examples to bootstrap step-by-step reasoning traces and filter out low-scoring demonstration candidates.