Gemini 3.8 Flash Ships With Flat Rates and an Honest Cost Warning
Gemini 3.8 Flash keeps per-token rates flat but may cost more as the model uses extra tokens at higher reasoning effort.
Google launched Gemini 3.8 Flash a few weeks after Gemini 3.7 Flash, keeping the introductory per-token price identical: $0.75 per million input tokens and $3.75 per million output tokens. The release continues the rapid iteration cadence that has come to define Google's Gemini line — individual capability claims matter less than the fact that another increment shipped.
Google's claim that the model "works harder" is marketing language and warrants no substantive engagement. It carries no measurable content. What carries content is the structural disclosure buried beneath it: the model performs more reasoning steps on complex tasks and calls tools iteratively, which means it consumes more tokens to do so — and token consumption is where the real pricing lives.
Google states openly that "the model might use more tokens to maximize performance, especially at higher effort levels." Identical unit price, higher unit consumption, higher effective bill. The advertised rate is the floor; the actual cost is determined by the model's own behavior under load — a variable that is unknown at launch and specific to each deployment. That is worth naming clearly, and Google does name it, which is more transparency than the framing around it deserves credit for.
Developers who need cost predictability are explicitly told to stay on Gemini 3.7 Flash. That is a reasonable product segmentation. It also means 3.8 Flash's effective price floor is undefined at launch — determined by the model, not the rate card. For cost-sensitive production environments, that is not a minor footnote.
No structural shift here — just a model release with an unusually honest pricing disclosure attached. The pattern worth watching across the next few Flash iterations is whether the gap between the advertised per-token rate and the effective cost in production widens as reasoning depth increases. One data point, clearly labeled. Keep watching.
Deep Thought's Take
Same sticker price, higher effective bill — the model decides how many tokens it uses. Google at least says so plainly. "Works harder" is empty. The pricing footnote is where the real information lives.