Gemini 3.8 Flash keeps its price and warns it will spend more tokens
Google's third Flash release in six weeks scores 73.7 percent on DeepSWE v1.1 at an unchanged $0.75 per million input tokens, and the blog says the model works harder.

Google shipped Gemini 3.8 Flash on September 2 at $0.75 per million input tokens and $3.75 per million output, the same rate as 3.7 Flash three weeks earlier. It scores 73.7 percent on DeepSWE v1.1, up from 65.3, and 54.9 percent on HLE-Verified. A second model, 3.8 Flash Cyber, went to a closed list of defenders through a new Fairwind Program.

Gemini 3.8 Flash on DeepSWE 1.1, scores 73.7%! t.co/JW7fhVy4He

The rate card did not move. Your bill probably will.
Here is the sentence to read twice, from Google's own post. "At times, the model might use more tokens to maximize performance, especially at higher effort levels." A page earlier the same post calls this a core design choice — 3.8 Flash exhibits greater diligence, executing extra reasoning steps and calling tools iteratively. Google then suggests that developers for whom compute efficiency is the primary constraint should turn the effort down, or stay on 3.7 Flash.
A per-token price that holds while per-task token consumption rises is a price increase written in a place the rate card cannot show. We have no numbers on how much more, because Google published none, and that is the gap worth watching.
The cadence is real and it is fast. July 21 for 3.6 Flash, August 13 for 3.7, September 2 for 3.8 — twenty-three days, then twenty. Across those forty-three days DeepSWE v1.1 went from 49.0 to 73.7, which is 24.7 points on a long-horizon software engineering benchmark in six weeks. Simon Willison had gemini-3.8-flash wired into his llm-gemini plugin within hours of the announcement (with low, medium and high thinking levels, which is where the token question lives).
Now the cyber model, where Google is more candid than the headline suggests. On CWE-Bench, an external patching benchmark run by Collinear, 3.8 Flash Cyber posts a pass@1 of 47.2 percent against a leading frontier model's 47.8. Google calls that being on the Pareto frontier (true, and also a loss). The internal numbers are better — a success rate above 70 percent finding vulnerabilities across twenty programming languages, 2.6 times more correct Chrome patches than much larger commercial models, and a critical vulnerability found by Google's Cloud Vulnerability Research team in under two hours where the work usually takes months.
Wiz, testing it on their own penetration-testing benchmark, reported 7.5 to 9.7 percent higher recall at 2.3 to 5.2 times lower cost. That is the shape of the whole release. Not better than the frontier. Cheaper at nearly the same place, which for most buyers is the better trade.
So how do you price a model that chooses how hard to work?
Badly, is our read, at least for anyone forecasting a budget. We would expect an independent cost-per-completed-task measurement to show 3.8 Flash costing more than 3.7 Flash on at least one agentic benchmark before the end of October, despite the identical rate. Google publishing median token counts per task would settle it, and would be the single most useful thing any lab could add to a model card right now.
And then there is the footnote. The $0.75 is introductory and expires on December 31. From January 1 it is $1.50 and $7.50 — double, on a model Google is telling you will also consume more tokens than the one it replaces.
