Review Machine · Master Review · AI libraries
DSPy 3.4.0
The strength its own pitch buries
Across fourteen sources read on 28 September 2026, three of the five dated write-ups describe DSPy 3.4.0's typed signatures as the contract a program is built on while the optimizers take the headlines, and one long developer thread raises the labelled set and metric a compile needs.
Version 3.4.0, released 25 September 2026Licence MITRepository 38,381 stars on 28 September 2026Read 28 September 2026Skip to the verdictRead as text

Verdict age
ProvisionalNo editorial review of 3.4 · developers thinThe official pages and the API call are three days younger than the 3.4.0 release; the developer thread is from March 2026 and the write-ups from April to September 2026. The project's 3.5 migration deadline would date the whole reading.
- Early · editors only
- Provisional · editors in, owners thin
- Settled · owner reviews read over months
From 14 sources to one verdict
Seven levels, in the same order in every Master Review. Each line is the takeaway; open a level for the evidence under it, or stop when you have enough.
What we read
14 sources, 5 Oct 2023 – 28 Sep 2026. Buyer, forum and other tiers thin. Editorial and news tiers empty.
By tier, heaviest first
The docs, the repository, the licence, the 3.4.0 release notes and the PyPI entry, read or dated within three days of the release.
One preprint, the project's own paper from October 2023. No established publication we could reach has reviewed the 3.4 line.
No reporting about this release was readable.
One REST API call on 28 September 2026: a figure on a date, not a quality signal.
One Hacker News thread from 23 March 2026, plus one search of its API on 28 September 2026.
Five write-ups, April to September 2026.
Date window
5 Oct 2023 to 28 Sep 2026
Editorial reviews 5 Oct 2023Official pages and the API call were read on 28 September 2026, three days after 3.4.0 shipped. The one developer thread is from 23 March 2026.
Couldn’t read, so not used
- DSPy Discord, 8.4k memberssign-in required
The community's Discord needs a signed-in account, so the developer tier here is Hacker News only.
Read, then rejected
- The New Stack, 10 July 2024 · describes the 2.x line, before every release named here
- Educative, 30 October 2025 · the 2.6 line; its no-optimizer result is about that line
What they measured
38,381 stars and 740 open issues on 28 September 2026, version 3.4.0 out 53 days after 3.3.0, and a compile the project reports at $2.18 over 200 examples.
Repository figures
38,381starson 28 September 2026
3,370forks
740open issues
GitHub · read 28 SepPackage
3.4.0versionreleased 25 September 2026
MITlicenceas the repository's licence file gives it
GitHub · read 28 Sep5.1M+downloads/mothe project's own figure on its home pageMaker's figure
DSPy · read 28 SepPyPI · read 28 SepRelease spacing
53days3.3.0 on 3 August 2026 to 3.4.0 on 25 September 2026
PyPI · read 28 SepA reported compile
$2.18costthe project's worked example: 200 labelled examples, 62% to 89%
DSPy · read 28 SepWhere they agree
The typed declaration is what the write-ups build a program on; the optimizers are what they lead with.
One tick per editorial review, in order of publication:AX arXiv
The signature is the contract
1/1The docs say signatures define tasks and enforce output types; three write-ups and the project's own paper describe the typed fields as the part a program rests on. T1Official T2Editorial Other
The compiled prompt is a file, not a string in the code
0/1The instructions are compiled into a JSON file loaded at runtime, so the prompt is an artifact. T1Official Other
The interface holds while the model changes
1/1The declaration stays fixed while the strategy or the model underneath is swapped. T1Official T2Editorial Forum Other
What keeps coming up against it
The optimizers want labelled examples and a metric, which one thread raises twice, and 3.x has moved the LM types.
Compiling wants a labelled set and a metric
Well documented · size unknownOne Hacker News thread raises it in two comments.
- Hacker News preparing training and test data is described as a reason the framework did not spread
- Hacker News the optimizers need datasets and tasks not every project has
Hacker News · 23 Mar
One dated write-up calls an undefined metric the common first mistake.
nolist.ai · read 28 Sep
The project's worked example compiles against 200 labelled examples.
DSPy · read 28 Sep
Compiled prompts are hard to read
0/1One write-up only, so it is not counted as recurring. Other
3.x moved the LM types
0/13.4 replaces the experimental types from 3.3 and 3.5 is the migration deadline; one write-up reports older tutorials breaking. T1Official Other
Extensions and the agent loop trail the alternatives
0/1Forum Other
A compile is a billable run
0/1T1Official Other
Where they split
Three places the evidence pulls two ways: on the optimizers, on the orchestration frameworks and on what to reach for.
Whether the optimizers are the point
One comment holds that the framework's real value is its prompt optimization, barely mentioned in the piece it answered.
Against that: A later comment takes the other side: the optimizers are not the point, and the contribution is pushing programmers to design a system that can be optimized.
Our readOne feature seen from two ends, the run you pay for and the structure it needs.
Whether it competes with the orchestration frameworks
sets DSPy against a large integration ecosystem and calls its optimizer the deeper engine.
Against that: says it does not orchestrate communicating agents the way those frameworks do, and calls the two categories complementary.
Our readA difference of framing rather than of fact.
What to reach for it for
recommends it to engineering teams building multi-stage retrieval pipelines that need systematic evaluation.
Against that: one commenter would take it now only for automating the edges of an agentic pipeline.
Our readBoth are single opinions.
Who it's for, who should pass
For a declared answer shape and a prompt you can save; pass if one call to one model is the whole job.
It suits you if
- You want the shape of a model's answer to be something you declared, not something you hope for.The docs say signatures define tasks and enforce output types, and the reply is parsed back into those fields.
- You have, or can label, a few dozen examples and a way to score answers.Every source here that describes compiling starts with examples and a metric.
- You expect to change the model under a working pipeline, or to keep the prompt as an artifact you can save.The docs make swapping the model a control, two write-ups describe the saved program, and one commenter describes a one-line model change followed by a rerun.
Pass if
- You need a single call to a single model and no evaluation.One dated write-up names simple wrapper applications as the case where the overhead outweighs the benefit.
- You need an interface that will not move between minor releases.The release notes replace the LM types in 3.4 with a deadline at 3.5, and one write-up reports older tutorials breaking on the move from 2 to 3.
- You want a capable agent loop out of the box.One write-up and one thread comment both point elsewhere for it.
The verdict
The optimizers are the headline; the typed declaration is what the write-ups and the thread keep using.
A framework whose typed declarations carry the work its optimizers get the credit for, read across five write-ups, one developer thread and the project's own pages.
Confidence, by tier
The docs, the repository, the licence, the 3.4.0 notes and PyPI, all read or dated within three days of the release.
One preprint from 2023; no established publication we could reach has reviewed the 3.4 line.
No reporting about this release was readable.
One REST API call: figures on a date, not a quality signal.
One Hacker News thread and one API search; themes, not ratings.
Five write-ups, none of them testing 3.4.0.
- Editorial evidence
- 35 months old
- Newest report
- None read
- Read
- 28 Sep · month 45
The compiler gets the applause. The typed signature does the work.
Rests onThree of the five dated write-ups we read describe the typed fields as the contract a program is built on, and the optimizers are named in the title or the summary of three of the same five.
Sources
14 sources, heaviest tier first. Every figure above comes from one of these.
- T1OfficialRelease notes 3.4.025 Sep
- T1OfficialDocs overviewread 28 Sep
- T1OfficialRepository and READMEread 28 Sep
- T1OfficialLicence file, MITread 28 Sep
- T1Officialdspy 3.4.0read 28 Sep
- T2EditorialDSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines5 Oct 2023
- T3BuyersRepository figures, one REST API callstars, forks and open issues, one callread 28 Sep
- ForumIf DSPy is so great, why isn't anyone using it?23 Mar
- ForumSite API search for DSPy12,181 stories matching, one callread 28 Sep
- OtherStop Writing Prompts, Start Optimizing Them27 Jun
- OtherDSPy, declaring not promptingno date printed; its address places it in April 2026read 28 Sep
- OtherDSPy Explainedno date printed on the pageread 28 Sep
- OtherDSPy Review 2026no date of its own; reports 3.3.1 as currentread 28 Sep
- OtherDSPy reviewno date; its getting-started section says March 2026read 28 Sep
MethodOn 28 September 2026 we read five of the project's own pages, one set of GitHub REST API figures, one preprint, two Hacker News items and a search, and five dated write-ups, by web search and plain page fetch, and synthesised them with AI; we ran nothing and tested nothing.
Dates without a year are 2026.
The review as text
DSPy is the framework from Stanford NLP that asks you to program a language model rather than prompt it, read here at version 3.4.0, three days after that release shipped on 25 September 2026.
The sample is fourteen sources read on 28 September 2026: five of the project's own pages and one set of repository figures, one preprint, two Hacker News items and one search of its API, and five write-ups dated between April and September 2026. The repository showed 38,381 stars, 3,370 forks and 740 open issues that day, and the project's home page claimed more than five million monthly downloads. No established publication we could reach has reviewed the 3.4 line, so the editorial tier is the project's own 2023 preprint and nothing more. We ran nothing: no model was prompted, no library installed and no benchmark rerun.
Consensus
The writing we could read agrees on what the framework is and divides on which half earns its place. The typed declaration is the half that turns up in passing: the project's docs say signatures define tasks and enforce output types, and three of the five dated write-ups we read describe those typed fields as the contract a program is built on. The optimizers are the half that turns up in the headline, named in the title or the summary of three of the same five write-ups.
Our confidence is strongest on the project's own pages, which are claims about themselves before they are evidence, and thinner on developer discussion: one thread from March 2026 and one API search. No editorial review of this line was readable, no news reporting either, and the buyer tier is a single API call. Two things could date this reading quickly: the 3.4.0 notes replace the LM types introduced in 3.3 and set 3.5 as the migration deadline, and 3.4.0 arrived fifty-three days after 3.3.0.
Recurring strengths
The signature carries the shape. The docs describe signatures as typed inputs and outputs that replace hand-managed prompts, and three of the five dated write-ups describe that layer as a contract rather than a convenience. One sets the round trip out as declaration, compiled prompt, then a reply parsed back into the typed fields it calls parse targets. One says the declaration is what makes a pipeline portable and testable. One lists typed signatures as the framework's first primitive, and the project's own 2023 paper sets out the same model in place of hard-coded prompt templates.
The compiled prompt is a file. Two of the five write-ups and the project's own pages mention the same deployment detail in passing: the instructions are compiled into a JSON file loaded at runtime. One advises saving the compiled state so an application does not rerun the optimization loop at every start. The 3.4.0 release notes carry a section on saved state, and the project's worked example ends by saving its program to a file.
The interface holds while the model changes. One write-up says the signature stays stable while the predictor strategy adapts around it, and another describes the modular design as separating program logic from model parameters. The project's home page makes the same point in a control that swaps the model, one commenter in the March thread describes trying one with a one-line change, and the project's 2023 paper reports compiled programs running on open and small models.
Recurring complaints
Compiling wants a labelled set and a metric. One Hacker News thread from March 2026 raises it in two comments: one describes preparing training and test data as a reason it did not spread, another says the optimizers need datasets and tasks not every project has. One dated write-up calls an undefined metric the common first mistake, since without one the optimizer has nothing to improve against. The project's worked example compiles against 200 labelled examples.
Compiled prompts are hard to read. One dated write-up lists debugging compiled prompts as abstract and difficult, and says the string sent to the model sits behind layers of abstraction. No other source raises it, so it stands as one complaint.
The 3.x line moved its LM types. The release notes call 3.4 the LM transition release and set 3.5 as the migration deadline, replacing the experimental types from 3.3 and warning that custom adapters must not rely on a removed module. One dated write-up says the move from version 2 to version 3 changed the module API enough that older tutorials throw validation errors, and one piece read in September 2026 still reports 3.3.1 as current.
Extensions and the agent story trail the alternatives. One write-up says the documentation is thinner than a mature competitor's, and one commenter says the out-of-the-box agent loop is the weak part.
A compile is a billable run. The DSPy docs' own worked example reports a compile over 200 examples costing $2.18 and taking its own figure from 62% to 89%, and one dated write-up treats avoiding a rerun of the optimization loop as a reason to save it.
Where reviewers split
Whether the optimizers are the point. One comment in the March thread holds that the framework's real value is its prompt optimization. A later comment takes the other side, that the optimizers are not the point and the contribution is pushing programmers to design a system that can be optimized at all. Our read: one feature seen from two ends, the run you pay for and the structure it needs.
Whether it competes with the orchestration frameworks. One dated write-up sets DSPy against a large integration ecosystem and calls its optimizer the deeper engine. Another says it does not orchestrate communicating agents the way those frameworks do, and calls the two categories complementary. We read this as a difference of framing rather than of fact.
What to reach for it for. One dated write-up recommends it to engineering teams building multi-stage retrieval pipelines that need systematic evaluation. In the thread, one commenter says the only thing they would take it for now is automating the edges of an agentic pipeline. Both are single opinions.
Who it suits
You want the shape of a model's answer to be something you declared, not something you hope for. The docs say signatures define tasks and enforce output types, and the framework parses the reply back into those fields.
You have, or can label, a few dozen examples and a way to score answers. Every source here that describes compiling, the project's own example included, starts with examples and a metric.
You expect to change the model under a working pipeline, or to keep the prompt as an artifact you can save. The docs make swapping the model a control, two write-ups and the release notes describe the saved program, and one commenter describes a one-line model change followed by a rerun.
Who should pass
You need a single call to a single model and no evaluation. One dated write-up names simple wrapper applications as the case where the architectural overhead outweighs the benefit.
You need an interface that will not move between minor releases. The release notes replace the LM types in 3.4 with a deadline at 3.5, and one write-up reports older tutorials breaking on the move from 2 to 3.
You want a capable agent loop out of the box. One write-up and one thread comment both point elsewhere for it.
Sources
- DSPy docs overview (official, read 28 September 2026)
- stanfordnlp/dspy repository (official, read 28 September 2026)
- DSPy licence file (official, MIT, read 28 September 2026)
- DSPy 3.4.0 release notes (official, published 25 September 2026)
- dspy 3.4.0 on PyPI (official, read 28 September 2026)
- GitHub repository figures (tier 3, read 28 September 2026)
- arXiv 2310.03714: DSPy, compiling declarative LM calls (editorial, published 5 October 2023)
- Hacker News: If DSPy is so great, why isn't anyone using it? (forum, published 23 March 2026)
- Hacker News API search for DSPy (forum, read 28 September 2026)
- Stories from a Software Tester: DSPy, declaring not prompting (other, April 2026 address)
- DataAspirant: DSPy Explained (other, read 28 September 2026)
- PromptQuorum: DSPy Review 2026 (other, read 28 September 2026)
- nolist.ai: DSPy review (other, read 28 September 2026)
- DevSkrol: Stop Writing Prompts, Start Optimizing Them (other, published 27 June 2026)
On 28 September 2026 we read five of the project's own pages, one set of GitHub REST API figures, one preprint, two Hacker News items and a search, and five dated write-ups, by web search and plain page fetch, and synthesised them with AI; we ran nothing and tested nothing.
One verdict a week.
Every week, the most useful Master Review we finished: what the internet agrees on, the complaint that kept appearing, and who should skip it.