Google shipped Gemini 3.8 Flash on 2 September 2026, its third release in this tier in six weeks, and every headline said the same thing: same price as 3.7, $0.75 per million input tokens and $3.75 output. Accurate, and it buries the part that will actually hit your invoice.
That $0.75 is an introductory rate with a hard expiry. Google's pricing page states it plainly for every current model in the tier: $0.75 through December 31, 2026, then $1.50 starting January 1, 2027. The launch announcement carries the same footnote. On 1 January your spend doubles overnight, on the same model, with no migration and no announcement to come.
The other half of the story is that 3.6 is on that discount too. Google quietly extended the introductory rate backwards across the whole line, so the older model you may have picked to save money now costs exactly what the newest one does.
Table of contents
- What Google's pricing page actually says
- The 2027 cliff, tier by tier
- Three generations, one price
- What 3.8 gives you for the same money
- Budgeting across the 1 January line
- FAQ
- Sources
What Google's pricing page actually says
The paid standard tier for gemini-3.8-flash, quoted from the pricing docs:
| Item | Through 31 Dec 2026 | From 1 Jan 2027 |
|---|---|---|
| Input / 1M tokens | $0.75 | $1.50 |
| Output / 1M (incl. thinking) | $3.75 | $7.50 |
| Context caching / 1M | $0.075 | $0.15 |
| Cache storage / 1M / hour | $0.50 | $1.00 |
Every line doubles. Not the headline rate with the rest held steady, the whole column. Cache storage included, which is the one people forget because it bills by time rather than by token.
Output is priced including thinking tokens, and Google says 3.8 Flash deliberately spends more of them. From the launch post: on complex tasks the model executes extra reasoning steps and calls tools iteratively, and "at times, the model might use more tokens to maximize performance, especially at higher effort levels." So your output bill can rise on identical prompts before the rate change touches it. Effort level is the lever if that matters to you.
The 2027 cliff, tier by tier
The doubling is uniform across serving tiers, which is worth seeing in one place because the batch and flex discount is where most people should be looking right now.
| Tier | Input now | Input 2027 | Output now | Output 2027 |
|---|---|---|---|---|
| Standard | $0.75 | $1.50 | $3.75 | $7.50 |
| Batch | $0.375 | $0.75 | $1.875 | $3.75 |
| Flex | $0.375 | $0.75 | $1.875 | $3.75 |
| Priority | $1.35 | $2.70 | $6.75 | $13.50 |
Batch and flex are a flat 50% off standard, and both keep that ratio after the change. Priority is 1.8x standard. Read the table across rather than down: batch pricing in 2027 ($0.75 / $3.75) is exactly today's standard pricing. If your workload tolerates asynchronous execution, moving it to batch cancels the increase entirely. That is the cleanest mitigation available and it needs no model change.
OpenRouter's live model API confirms the same numbers from outside Google: google/gemini-3.8-flash at $0.75 / $3.75, and google/gemini-3.8-flash:batch at $0.375 / $1.875, with a 1,048,576-token context on both.
Three generations, one price
Here is the part I did not expect. I checked the pricing page for each generation still listed, expecting a ladder.
| Model | Input | Output | Cached input |
|---|---|---|---|
| Gemini 3.8 Flash | $0.75 | $3.75 | $0.075 |
| Gemini 3.7 Flash | $0.75 | $3.75 | $0.075 |
| Gemini 3.6 Flash | $0.75 | $3.75 | $0.075 |
Identical, and all three expire on the same date. Google's own page calls 3.6 "our previous generation" while charging the same rate as the current one.
If you pinned gemini-3.6-flash in production for cost reasons, that reason no longer exists. You are paying current-generation prices for a model Google has now superseded twice. Pinning for output stability is still a real argument, and reasoning models do shift behaviour between versions in ways that break brittle prompts. Just make that the stated reason instead of a cost saving you are not receiving.
Our API cost calculator now carries all three at the correct $0.75 / $3.75, so you can put your own token mix against them instead of trusting a table you read somewhere. It previously listed 3.6 at the post-2027 rate, which overstated the cost by 2x.
What 3.8 gives you for the same money
Google's claimed gains over 3.7, from the launch post:
- 54.9% on HLE-Verified for multi-step reasoning across STEM, humanities and professional fields
- Beats most larger frontier models on DeepSWE v1.1 for long-horizon software engineering, at a fraction of their cost
- Ahead of 3.7 and other frontier models on Vals Finance Agent V2 and Harvey's Legal Agent Benchmark
These are vendor benchmarks, so treat them as a claim rather than a measurement. The structural point holds regardless: when a newer model costs the same as the one it replaces, there is no cost case for staying behind, only a stability case.
The companion release is Gemini 3.8 Flash Cyber, gated to trusted defenders through Google's new Fairwind Program. It has no public price because you cannot buy it. Google reports the Chrome Security team getting 2.6x more correct Chrome vulnerability patches from it than from much larger commercial models, and 47.2% pass@1 on CWE-Bench against a leading frontier model's 47.8%. Both variants run on the same foundational model, and Google credits cybersecurity training for some of the coding gains in the public one.
Budgeting across the 1 January line
What I would do with a Flash-heavy bill, in order.
Model both prices in whatever spreadsheet or dashboard your finance team reads. A Q1 2027 forecast built on $0.75 is wrong by 100% on the token line, and that is the kind of error that surfaces in February when someone asks why the bill moved.
Move every asynchronous job to batch before the deadline rather than after. Evaluation runs, backfills, bulk classification, nightly summarisation: batch keeps them at today's effective rate through 2027. Doing it now also means the migration is not competing with a January cost panic.
Stop paying a premium for an older version. There is no premium, which means there is also no discount. Pick the version on behaviour and stability, not price.
Watch context caching if you run long agentic sessions. The per-token cache rate doubles and so does hourly storage, so a session that holds a large cache for hours takes the increase twice. Caching still wins against re-sending tokens at $1.50, but the margin narrows.
Do not assume a fourth release resets the clock. Three of them in six weeks makes it likely another lands before January, but 3.8 inherited 3.7's expiry date rather than getting a fresh twelve months. Ars Technica's read is that new models will arrive long before the price changes, which is probably true and does not help. The expiry is attached to the calendar, not to the model.
FAQ
How much does Gemini 3.8 Flash cost? $0.75 per million input tokens and $3.75 per million output tokens on the paid standard tier through 31 December 2026. From 1 January 2027 it is $1.50 and $7.50.
Is that price permanent? No. Google's pricing page and the launch announcement both label it introductory, with an explicit expiry of 31 December 2026 and a stated post-expiry rate of $1.50 / $7.50.
Does Gemini 3.6 Flash cost less than 3.8? No. Gemini 3.6, 3.7 and 3.8 are all $0.75 input and $3.75 output, with the same 31 December 2026 expiry.
What is the cheapest way to run it? Batch or flex, both at $0.375 input and $1.875 output, half the standard rate. Their 2027 prices match today's standard prices, so batch workloads absorb the increase.
Do thinking tokens count as output? Yes. Google prices output "including thinking tokens," and says the model may spend more of them on complex tasks at higher effort levels. Lower the effort level to cut token overhead.
What is Gemini 3.8 Flash Cyber and what does it cost? A cybersecurity variant restricted to trusted defenders through Google's Fairwind Program. No public pricing, since access is by application rather than by API key.
What is the context window? 1,048,576 tokens, per OpenRouter's model listing for both the standard and batch endpoints.
Conclusion
Gemini 3.8 Flash at $0.75 and $3.75 is a good rate for a model this capable, and it is worth switching to from 3.6 or 3.7 because those cost the same. Take the free upgrade.
The number to write down is 1 January 2027, when every line of that table doubles. If a meaningful share of your Flash traffic can run asynchronously, batch it before then and the increase never reaches you.
Sources
- Gemini Developer API pricing — official tier tables and expiry dates, Wayback snapshot 6 September 2026
- Introducing Gemini 3.8 Flash and 3.8 Flash Cyber — launch benchmarks, effort-level note and the introductory-price footnote, read 7 September 2026
- OpenRouter models API — third-party confirmation of standard and batch rates plus context length, read 7 September 2026
- Google releases Gemini 3.8 Flash, its third Flash model in six weeks — release cadence, 2 September 2026
