Review Machine · Master Review · ai-models
Gemini 3.8 Flash (gemini-3.8-flash)
The complaints that follow Google's newest Flash model, counted across 24 sources
Across 24 sources read on 27 September 2026, Gemini 3.8 Flash is a post-trained 3.7 Flash at the top of the long-horizon coding boards, with four developer sources reporting the waiting and five keeping an older Flash instead.
Model id gemini-3.8-flash, stableReleased 2 September 2026, generally availableBuilt on Gemini 3.7 Flash, per the model cardWindow 1,048,576 in and 65,536 outRead 27 September 2026Skip to the verdictRead as text

Verdict age
ProvisionalDevelopers in, editors thinThe model is 25 days old at the day of reading. The coding boards were updated within a day of launch and the developer threads run to 25 September. The rate card is the firmest thing here, and it changes on 1 January 2027.
- Early · editors only
- Provisional · editors in, owners thin
- Settled · owner reviews read over months
From 24 sources to one verdict
Seven levels, in the same order in every Master Review. Each line is the takeaway; open a level for the evidence under it, or stop when you have enough.
What we read
24 sources, 2 Sep – 27 Sep 2026. Editorial and news tiers thin. Buyer tier empty.
By tier, heaviest first
Eight pages Google publishes, dated 2 to 25 September 2026.
Independent evaluations rather than publication reviews: two dated benchmark boards and one tracker with published intervals.
Launch-day reporting. No publication we could read had reviewed the model itself.
No buyer tier exists for a model. Nobody rates an API the way owners rate a phone.
Eleven developer sources dated 2 to 25 September 2026, the loudest and most numerous tier here.
Date window
2 Sep 2026 to 27 Sep 2026
Editorial reviews 2 Sep – 3 Sep 2026 · News 2 Sep 2026The model is 25 days old. Everything read is dated between its launch and the day of reading.
Couldn’t read, so not used
- Ars Technica report on the releaseHTTP 405 on a plain page fetch
- The Verge report on the releasebody would not render, headline only
- Lifehacker hands-onbody would not render, headline only
- Reddit r/GeminiAI and r/google_antigravity threadsHTTP 403, sign-in wall
Nothing above is cited, and no claim rests on any of them.
Read, then rejected
- Business Insider and Wall Street Journal syndication on the pre-release build · a leak before launch, not evidence about the released model
- lmmarketcap, opencode.ai, requesty, modelcap and other aggregators · reseller or comparison pages rather than evaluations
- Blogs restating Google's rate card · Google's pricing page was read directly
- LiteLLM day-zero support post · vendor note; its one limitation is restated from Google's own pages
What they measured
The independent figures we could read put this model level with the top entry on the long-horizon coding board, and its ceilings are unchanged from 3.7 Flash.
Long-horizon coding, DeepSWE v1.1 at high effort
74% ±1level with the board's top entry, a one-point interval
DeepSWE · updated 3 Sep73.7%Google's own claimed figure on the same board, against 65.3% for 3.7 Flash
Google DeepMind · 2 SepCost of one Intelligence Index task at high effort
$0.58per taskThe same index puts Anthropic's Fable 5.1 at $3.76 a task
The Register · 2 SepRate card per million tokens
$0.75in / 1Mthrough 31 December 2026, then $1.50
Google · read 27 Sep$3.75out / 1Mincludes thinking tokens; $7.50 from 1 January 2027
Google · read 27 Sep$0.075cache /1Mthen $0.15; storage $0.50 per million tokens an hour
Google · read 27 SepCeilings
1,048,576tokens in
Google · read 27 Sep65,536tokens out
Google · read 27 SepSelf-initiated tool calls in one developer's comparison
0of 8the same prompt against 3.5 Flash scored 8 of 8; 6 of 8 with an explicit rule addedOne developer's test
Google AI Developers Forum · 5 SepFree tier, requests a day
20requestsFlash-Lite allows 500 a dayForum report
Google AI Developers Forum · 3 SepWhere they agree
Two independent evaluations put it at the top of the long-horizon coding boards at high effort, and every source we read agrees on what the model is: a post-training pass on 3.7 Flash at the same rate card.
One tick per editorial review, in order of publication:AA Artificial AnalysisDS DeepSWETM The Model Gap
Top of the long-horizon coding boards at high effort
1/3The DeepSWE v1.1 board lists it level with its best entry; Google's pages put DeepSWE v1.1 at 73.7% and Terminal-Bench 2.1 between 89.4% and 90.8%. T1Official T2Editorial
The same envelope as 3.7 Flash, at the same rate card
1/3Post-training rather than a new pretraining run, on the same window, the same output ceiling and the same per-token price until 31 December 2026. T1Official T2Editorial T2News
What keeps coming up against it
The complaints are about waiting and about tokens: four developer sources report latency or capacity, four report quota burned by the model's own reading, and five keep an older Flash instead.
Slowness and capacity in the tools built on it
Well documented, size unknownThe model card lists slowness or timeout issues among the model's known limitations, and says it may use more tokens to maximise performance at higher effort levels.
Google DeepMind · 2 Sep
Four threads dated 15 to 25 September 2026: a median step turnaround of 16.0 seconds against 3.0 seconds two days earlier; 503s saying no capacity is available between 16:00 and 22:00 UTC with a two-minute task past twenty minutes; latency at every thinking level; and a thread titled in the developer's own words.
- Google AI Developers Forum 16.0 seconds a step where two days earlier it was 3.0
- Google AI Developers Forum 503, no capacity for the high-effort model, between 16:00 and 22:00 UTC
- Google AI Developers Forum Latency at every thinking level, with quota billing for the failures
Google AI Developers Forum · 15 SepGoogle AI Developers Forum · 15 SepGoogle AI Developers Forum · 18 SepGoogle AI Developers Forum · 25 Sep
None of the three launch-day reports we read puts a figure on latency; the complaints are all in developer threads, dated from 3 September onward.
Waiting, and 503s at peak hours
0/3Four developer threads: 16.0 seconds a step against 3.0 two days earlier, capacity errors between 16:00 and 22:00 UTC, latency at every thinking level, and a thread titled too slow. Google's card lists slowness among known limitations. T1Official Forum
Tokens spent and quota burned
1/3The launch thread counts 120 million tokens on the index at high effort against 64 million for 3.7 Flash; one thread reports full-file reads draining an Ultra-tier quota and another keeps 3.6 for file-heavy work; Artificial Analysis calls the model very verbose against its 88-million-token median. T2Editorial Forum
Rolling back to an earlier Flash
0/3Five developer sources: a huge regression with reasoning called 2.5-level, a switch back to 3.7 the same day, self-initiated tool calls at 0 of 8 against 3.5 Flash's 8 of 8, and two threads keeping 3.6 or 3.7 for file-heavy work. Forum
Loops and retries
0/3Repeated shell-command loops across seven sessions in one evening, and getting stuck re-reading the same files instead of fixing the bug. Forum
Parts of the 3.7 interface removed
0/3The level that turns thinking off now errors, temperature, top_p and top_k are ignored, penalties and a candidate count throw, prefilled model turns must go, and the Live API is not supported here. T1Official
A free tier too small to evaluate on
0/3One developer thread argues twenty requests a day cannot evaluate a model in an agent workflow, against five hundred for Flash-Lite. Forum
Every rate doubles on 1 January 2027
0/3Input to $1.50 and output to $7.50 per million tokens, batch and flex to $0.75 and $3.75, cache reads to $0.15; the free tier is the only line that does not move. T1Official T2News
Where they split
The maker's own two pages disagree on the same benchmark, the gains over 3.7 Flash sit inside the benchmarks' noise bands, and the boards point the opposite way from the developers.
Google's own pages, on the same benchmark
An HLE-Verified score of 54.9% and Terminal-Bench 2.1 at 89.4%.
Against that: HLE at 45.4% and Terminal-Bench 2.1 at 90.8% for the same model.
Both publishedTwo pages Google publishes give different figures for the same benchmark. We report both and pick neither.
How much the gain over 3.7 Flash is worth
Level with 3.7 Flash on Humanity's Last Exam at 47.8 against 47.9 and fractionally behind on SWE-bench Verified at 80.0 against 80.8, each difference inside the benchmark's own noise band.
Against that: DeepSWE v1.1 up from 65.3% to 73.7% and a claimed HLE-Verified figure with no interval published beside it.
The boards against the developers
Both put this model above 3.7 Flash on the boards they measure.
Against that: Report the other direction and keep 3.6 or 3.7 for their own work.
The Flash model and the missing Pro
Google's own comparison figures still put this model behind Claude Opus 5 at agentic computer use, and the Cyber variant released beside it is limited to trusted testers and unnamed governments.
Against that: The launch post presents it as the best reasoning and coding model Google has shipped, at the same introductory price as 3.7 Flash.
Who it's for, who should pass
It fits long-horizon agentic work where the board position is the point and the bill is per task; it does not fit work that needs a predictable wait, a stable per-token bill, or the parts of the old interface.
It suits you if
- You are building long-horizon coding or agentic runs.Two independent evaluations put it level with the top entry on the DeepSWE v1.1 board at high effort (editorial, 3 September 2026).
- You want the newest Flash at the rate card you already budgeted against.The price per token is unchanged from 3.7 Flash until 31 December 2026 (official, read 27 September 2026).
- You need one model carrying a million-token window, thinking levels, tool calls, grounding and a batch tier.All of them are supported on this model id (official, read 27 September 2026).
- You can route latency-sensitive steps to an older Flash.Five of the eleven developer sources we read did exactly that (forum, 3 to 20 September 2026).
Pass if
- You need a predictable wait at a predictable hour.One developer reports 503s between 16:00 and 22:00 UTC and a two-minute task past twenty minutes, and three more report high waits (forum, 15 to 25 September 2026).
- You pay per token on a high-volume pipeline.The reading counts the model spending more tokens a task, and every rate doubles on 1 January 2027 (editorial and official, 2 September 2026 and 27 September 2026).
- You depend on the interface parts that were removed, or on the Live API, tuning or a larger free tier.The level that turns thinking off errors, sampling parameters are ignored, and the Live API is not supported here (official, 23 and 25 September 2026).
The verdict
You can have the top of the long-horizon coding boards at an unchanged rate card, and you will spend more tokens and wait longer for it.
Gemini 3.8 Flash sits level with the best long-horizon coding result we could read, while four developer sources report the waiting and five keep an older Flash instead.
Confidence, by tier
Eight pages read directly, dated 2 to 25 September 2026, including the price and the limits.
Two dated boards and one tracker with published intervals; no publication we could read reviewed the model itself.
Two launch-day reports, neither revisiting the model later.
No buyer platform rates a model. The tier does not exist for this subject.
Eleven developer sources dated 3 to 25 September 2026, consistent with each other on latency, tokens and rollbacks.
- Editorial evidence
- under a month old
- Newest report
- 2 Sep 2026
- Read
- 27 Sep · month 0
Top of the long-horizon coding board. It reads your whole repository to get there.
Rests onDeepSWE v1.1 lists gemini-3.8-flash at high effort at 74% with a one-point interval, level with the board's best entry (editorial, 3 September 2026), and two developer threads dated 18 and 20 September 2026 describe full-file reads eating the quota, with Artificial Analysis calling the model very verbose against its 88-million-token median (editorial, 2 September 2026).
Sources
24 sources, heaviest tier first. Every figure above comes from one of these.
- T1OfficialGemini 3.8 Flash model card2 Sep
- T1OfficialRate limits, Gemini APIupdated 2 Sep
- T1OfficialIntroducing Gemini 3.8 Flash and 3.8 Flash Cyber2 Sep
- T1OfficialWhat is new in Gemini 3.8 Flashupdated 23 Sep
- T1OfficialGemini 3.8 Flash, Gemini Enterprise Agent Platformupdated 25 Sep
- T1OfficialDeveloper's guide to Gemini 3.8 Flashupdated 25 Sep
- T1OfficialGemini 3.8 Flash, Gemini API model pageread 27 Sep
- T1OfficialGemini Developer API pricingread 27 Sep
- T2EditorialGemini 3.8 Flash, intelligence, speed and cost2 Sep
- T2EditorialDeepSWE v1.1 leaderboardupdated 3 Sep
- T2EditorialGemini 3.8 Flash benchmarks and pricingread 3 Sep
- T2NewsWith Gemini 3.8 Flash, Google reminds everyone it is still in the race2 Sep
- T2NewsGoogle releases Gemini 3.8 Flash and a cybersecurity variant limited to governments2 Sep
- ForumGemini 3.8 Flash and 3.8 Flash Cyber, launch thread2 Sep
- ForumIs Gemini 3.8 a huge regression3 Sep
- ForumGemini 3.8 Flash free tier 20 RPD is too limited for practical evaluation3 Sep
- ForumTool calls suppressed when the prompt contains reflective first-person text5 Sep
- ForumGemini 3.8 Live and 3.8 Live Extended Thinking15 Sep
- ForumGemini 3.8 in Antigravity is too slow15 Sep
- ForumSevere latency regression and time-to-first-token spike on Gemini 3.8 Flash15 Sep
- ForumSevere latency, command looping, quota drain and API 503 billing issues18 Sep
- ForumAntigravity Gemini models get stuck on reading repeated files nonstop20 Sep
- ForumGemini 3.6 Flash is better than 3.7 and 3.820 Sep
- ForumSevere slowdowns and frequent 503 errors on Gemini Flash 3.8 during specific hours25 Sep
MethodMethod: 24 sources dated 2 to 25 Sep 2026, read on 27 Sep 2026 and synthesised with AI; no model was prompted, no benchmark rerun and nothing tested.
Dates without a year are 2026.
The review as text
Gemini 3.8 Flash is Google's newest Flash-tier model, generally available on 2 September 2026 as gemini-3.8-flash and built, its model card says, on Gemini 3.7 Flash. We read 24 sources on 27 September 2026: eight of Google's pages, three independent evaluations, two launch-day reports and eleven developer sources, the Hacker News launch thread and nine threads on Google's AI Developers Forum among them. No buyer tier exists for a model, and no publication we could read had reviewed it beyond a first look, so this reading is official pages, independent boards and developers. What follows it is a complaint about what it costs to find out and how long it makes you wait.
Consensus
Across the sources we read, Gemini 3.8 Flash is the strongest Flash-tier model Google ships and not a new model at all: 3.7 Flash with more post-training, on the same 1,048,576-token input window, the same 65,536-token output ceiling and the same rate card. Two independent evaluations put it at the top of the long-horizon coding boards at high effort, and Google's model card claims 73.7% on the same board against 65.3% for 3.7 Flash. The confidence is real and narrow: the coding result is measured independently, little else is, and Google's own two pages disagree on the same benchmark.
Recurring strengths
Top of the long-horizon coding boards at high effort. DeepSWE's v1.1 board lists it at 74% with a one-point interval, level with the board's best entry (editorial, 3 Sep 2026). Google's pages put DeepSWE v1.1 at 73.7% and Terminal-Bench 2.1 between 89.4% and 90.8% (official, 2 and 25 Sep 2026).
Cheap per task at its level. The Register reports $0.58 per Intelligence Index task against $3.76 for the same index on Anthropic's Fable 5.1 (news, 2 Sep 2026). The rate is unchanged from 3.7 Flash: $0.75 per million input tokens and $3.75 per million output (official, 27 Sep 2026).
The envelope carries over. The window, the ceiling, thinking levels, caching, function calling, grounding and the batch, flex and priority tiers all carry over from 3.7 Flash (official, 27 Sep 2026).
Recurring complaints
Waiting, and 503s at peak hours. Four of the eleven developer sources describe this, and it is the loudest complaint in the reading. One thread measures a median step turnaround of 16.0 seconds at high effort against 3.0 seconds two days earlier; another reports 503s saying no capacity is available between 16:00 and 22:00 UTC, a two-minute task past twenty minutes and speed down to a tenth; a third reports high latency at every thinking level and a fourth is titled that the model is too slow (forum, 15 to 25 Sep 2026). Google's model card lists slowness and timeout issues among its known limitations (official, 2 Sep 2026).
Tokens and quota. Four developer sources and one of the three evaluations. The launch thread counts 120 million tokens spent on the index at high effort against 64 million for 3.7 Flash (forum, 2 Sep 2026). One thread reports full-file reads draining an Ultra-tier quota, and another keeps 3.6 Flash for file-heavy work (18 and 20 Sep 2026). Artificial Analysis calls the model very verbose against its median of 88 million output tokens (editorial, 2 Sep 2026).
Rolling back to an earlier Flash. Five of the eleven. One thread calls the model a huge regression, placing its reasoning at 2.5-level on a trivial question (3 Sep 2026); the Live thread reports a switch back to 3.7 (15 Sep 2026); a third measures self-initiated tool calls at 0 of 8 where 3.5 Flash scores 8 of 8 (5 Sep 2026); two more keep 3.6 or 3.7 and call 3.6 the better balanced model (20 Sep 2026).
Loops and retries. Two of the eleven. Repeated shell-command loops across seven sessions in one evening, and getting stuck re-reading the same files instead of fixing the bug (18 and 20 Sep 2026).
Parts of the 3.7 interface are gone. The level that turns thinking off errors, temperature, top_p and top_k are ignored, penalties and a candidate count throw, prefilled turns must be removed, and the Live API is unsupported (official, 23 and 25 Sep 2026).
A free tier too small to evaluate on. One thread argues twenty requests a day cannot evaluate a model in an agent workflow, against five hundred for Flash-Lite (3 Sep 2026).
Every rate doubles on 1 January 2027. Input to $1.50 and output to $7.50 per million tokens, batch and flex to $0.75 and $3.75, cache reads from $0.075 to $0.15; the free tier is the only line that does not move (official, 27 Sep 2026).
Where reviewers split
Google's own pages disagree. The launch post and the model card give HLE-Verified at 54.9%; the Gemini Enterprise developer guide gives HLE at 45.4% for the same model, and the same pair put Terminal-Bench 2.1 at 89.4% and 90.8% (official, 27 Sep 2026).
The gain over 3.7 Flash sits inside the benchmarks' own noise. The Model Gap, which tracks boards with published intervals, records it level with 3.7 Flash on Humanity's Last Exam at 47.8 against 47.9 and fractionally behind on SWE-bench Verified at 80.0 against 80.8 (editorial, 3 Sep 2026).
The boards and the developers point opposite ways. Both evaluations we read put this model above 3.7 Flash; five of the eleven developer sources report the other direction and keep the older model (editorial and forum, 2 to 25 Sep 2026).
And outside the model, the missing Pro. The Next Web notes that Google's own figures still put this Flash behind Claude Opus 5 at agentic computer use, and that the Cyber variant beside it goes only to trusted testers and unnamed governments (news, 2 Sep 2026).
Who it suits
You are building long-horizon coding or agentic runs. Two independent evaluations put it level with the top entry on the DeepSWE v1.1 board at high effort (editorial, 3 Sep 2026).
You want the newest Flash at the rate you already budgeted against, and can finish before 1 January 2027 (official, 27 Sep 2026).
You need one model carrying a million-token window, thinking levels, tool calls, grounding and a batch tier (official, 27 Sep 2026).
Who should pass
You need a predictable wait at a predictable hour. One developer reports 503s between 16:00 and 22:00 UTC and a two-minute task past twenty minutes, and three more report high waits (forum, 15 to 25 Sep 2026).
You pay per token on a high-volume pipeline. The reading says the model spends more tokens a task, and every rate doubles in January (editorial and official, 2 Sep 2026 and 27 Sep 2026).
You depend on the parts of the 3.7 interface that were removed, or need the Live API, tuning, or a free tier you can evaluate thirty agent calls on (official, 23 and 25 Sep 2026).
One verdict a week: the most useful Master Review we finished, the complaint that kept appearing, and who should skip it. The weekly verdict is here.
Sources
- Model card, DeepMind, official, 2 Sep 2026
- Gemini API model page, Google, official, 27 Sep 2026
- What is new in Gemini 3.8 Flash, Google, official, 23 Sep 2026
- Gemini Developer API pricing, Google, official, 27 Sep 2026
- Rate limits, Google, official, 2 Sep 2026
- Gemini Enterprise Agent Platform model page, Google Cloud, official, 25 Sep 2026
- Developer's guide to Gemini 3.8 Flash, Google Cloud, official, 25 Sep 2026
- Launch post, DeepMind, official, 2 Sep 2026
- Gemini 3.8 Flash, intelligence, speed and cost, Artificial Analysis, editorial, 2 Sep 2026
- DeepSWE v1.1 leaderboard, DeepSWE, editorial, 3 Sep 2026
- Gemini 3.8 Flash benchmarks and pricing, The Model Gap, editorial, 3 Sep 2026
- Google reminds everyone it is still in the race, The Register, news, 2 Sep 2026
- A cybersecurity variant limited to governments, The Next Web, news, 2 Sep 2026
- Launch thread, Hacker News, forum, 2 Sep 2026
- Gemini 3.8 Live and Extended Thinking, Hacker News, forum, 15 Sep 2026
- Is Gemini 3.8 a huge regression, dev forum, 3 Sep 2026
- Free tier 20 RPD is too limited, dev forum, 3 Sep 2026
- Tool calls suppressed by reflective first-person text, dev forum, 5 Sep 2026
- Gemini 3.8 in Antigravity is too slow, dev forum, 15 Sep 2026
- Time-to-first-token spike on Gemini 3.8 Flash, dev forum, 15 Sep 2026
- Latency, command looping, quota drain and API 503 billing, dev forum, 18 Sep 2026
- Stuck on reading repeated files nonstop, dev forum, 20 Sep 2026
- Gemini 3.6 Flash is better than 3.7 and 3.8, dev forum, 20 Sep 2026
- Severe slowdowns and 503 errors during specific hours, dev forum, 25 Sep 2026
Method: 24 sources dated 2 to 25 Sep 2026, read on 27 Sep 2026 and synthesised with AI; no model was prompted, no benchmark rerun and nothing tested.
One verdict a week.
Every week, the most useful Master Review we finished: what the internet agrees on, the complaint that kept appearing, and who should skip it.