Editorial

Gemini 3.8 Flash Pricing: The $0.75 Rate Expires Dec 31

Gemini 3.8 Flash is $0.75/$3.75 per MTok only until December 31, 2026. The tier table, the 2027 doubling, and how batch avoids it.

JJyoti Ranjan SwainUpdated
Gemini 3.8 Flash pricing: $0.75 and $3.75 per million tokens until 31 December 2026, then $1.50 and $7.50

Google shipped Gemini 3.8 Flash on 2 September 2026, its third release in this tier in six weeks, and every headline said the same thing: same price as 3.7, $0.75 per million input tokens and $3.75 output. Accurate, and it buries the part that will actually hit your invoice.

That $0.75 is an introductory rate with a hard expiry. Google's pricing page states it plainly for every current model in the tier: $0.75 through December 31, 2026, then $1.50 starting January 1, 2027. The launch announcement carries the same footnote. On 1 January your spend doubles overnight, on the same model, with no migration and no announcement to come.

The other half of the story is that 3.6 is on that discount too. Google quietly extended the introductory rate backwards across the whole line, so the older model you may have picked to save money now costs exactly what the newest one does.

Table of contents

What Google's pricing page actually says

The paid standard tier for gemini-3.8-flash, quoted from the pricing docs:

ItemThrough 31 Dec 2026From 1 Jan 2027
Input / 1M tokens$0.75$1.50
Output / 1M (incl. thinking)$3.75$7.50
Context caching / 1M$0.075$0.15
Cache storage / 1M / hour$0.50$1.00

Every line doubles. Not the headline rate with the rest held steady, the whole column. Cache storage included, which is the one people forget because it bills by time rather than by token.

Gemini 3.8 Flash standard-tier pricing doubles across every line item on 1 January 2027

Output is priced including thinking tokens, and Google says 3.8 Flash deliberately spends more of them. From the launch post: on complex tasks the model executes extra reasoning steps and calls tools iteratively, and "at times, the model might use more tokens to maximize performance, especially at higher effort levels." So your output bill can rise on identical prompts before the rate change touches it. Effort level is the lever if that matters to you.

The 2027 cliff, tier by tier

The doubling is uniform across serving tiers, which is worth seeing in one place because the batch and flex discount is where most people should be looking right now.

TierInput nowInput 2027Output nowOutput 2027
Standard$0.75$1.50$3.75$7.50
Batch$0.375$0.75$1.875$3.75
Flex$0.375$0.75$1.875$3.75
Priority$1.35$2.70$6.75$13.50

Batch and flex are a flat 50% off standard, and both keep that ratio after the change. Priority is 1.8x standard. Read the table across rather than down: batch pricing in 2027 ($0.75 / $3.75) is exactly today's standard pricing. If your workload tolerates asynchronous execution, moving it to batch cancels the increase entirely. That is the cleanest mitigation available and it needs no model change.

OpenRouter's live model API confirms the same numbers from outside Google: google/gemini-3.8-flash at $0.75 / $3.75, and google/gemini-3.8-flash:batch at $0.375 / $1.875, with a 1,048,576-token context on both.

Three generations, one price

Here is the part I did not expect. I checked the pricing page for each generation still listed, expecting a ladder.

ModelInputOutputCached input
Gemini 3.8 Flash$0.75$3.75$0.075
Gemini 3.7 Flash$0.75$3.75$0.075
Gemini 3.6 Flash$0.75$3.75$0.075

Identical, and all three expire on the same date. Google's own page calls 3.6 "our previous generation" while charging the same rate as the current one.

Gemini 3.6, 3.7 and 3.8 Flash all price at 0.75 dollars input and 3.75 output, expiring together on 31 December 2026

If you pinned gemini-3.6-flash in production for cost reasons, that reason no longer exists. You are paying current-generation prices for a model Google has now superseded twice. Pinning for output stability is still a real argument, and reasoning models do shift behaviour between versions in ways that break brittle prompts. Just make that the stated reason instead of a cost saving you are not receiving.

Our API cost calculator now carries all three at the correct $0.75 / $3.75, so you can put your own token mix against them instead of trusting a table you read somewhere. It previously listed 3.6 at the post-2027 rate, which overstated the cost by 2x.

What 3.8 gives you for the same money

Google's claimed gains over 3.7, from the launch post:

  • 54.9% on HLE-Verified for multi-step reasoning across STEM, humanities and professional fields
  • Beats most larger frontier models on DeepSWE v1.1 for long-horizon software engineering, at a fraction of their cost
  • Ahead of 3.7 and other frontier models on Vals Finance Agent V2 and Harvey's Legal Agent Benchmark

These are vendor benchmarks, so treat them as a claim rather than a measurement. The structural point holds regardless: when a newer model costs the same as the one it replaces, there is no cost case for staying behind, only a stability case.

The companion release is Gemini 3.8 Flash Cyber, gated to trusted defenders through Google's new Fairwind Program. It has no public price because you cannot buy it. Google reports the Chrome Security team getting 2.6x more correct Chrome vulnerability patches from it than from much larger commercial models, and 47.2% pass@1 on CWE-Bench against a leading frontier model's 47.8%. Both variants run on the same foundational model, and Google credits cybersecurity training for some of the coding gains in the public one.

Budgeting across the 1 January line

What I would do with a Flash-heavy bill, in order.

Model both prices in whatever spreadsheet or dashboard your finance team reads. A Q1 2027 forecast built on $0.75 is wrong by 100% on the token line, and that is the kind of error that surfaces in February when someone asks why the bill moved.

Move every asynchronous job to batch before the deadline rather than after. Evaluation runs, backfills, bulk classification, nightly summarisation: batch keeps them at today's effective rate through 2027. Doing it now also means the migration is not competing with a January cost panic.

Stop paying a premium for an older version. There is no premium, which means there is also no discount. Pick the version on behaviour and stability, not price.

Watch context caching if you run long agentic sessions. The per-token cache rate doubles and so does hourly storage, so a session that holds a large cache for hours takes the increase twice. Caching still wins against re-sending tokens at $1.50, but the margin narrows.

Do not assume a fourth release resets the clock. Three of them in six weeks makes it likely another lands before January, but 3.8 inherited 3.7's expiry date rather than getting a fresh twelve months. Ars Technica's read is that new models will arrive long before the price changes, which is probably true and does not help. The expiry is attached to the calendar, not to the model.

FAQ

How much does Gemini 3.8 Flash cost? $0.75 per million input tokens and $3.75 per million output tokens on the paid standard tier through 31 December 2026. From 1 January 2027 it is $1.50 and $7.50.

Is that price permanent? No. Google's pricing page and the launch announcement both label it introductory, with an explicit expiry of 31 December 2026 and a stated post-expiry rate of $1.50 / $7.50.

Does Gemini 3.6 Flash cost less than 3.8? No. Gemini 3.6, 3.7 and 3.8 are all $0.75 input and $3.75 output, with the same 31 December 2026 expiry.

What is the cheapest way to run it? Batch or flex, both at $0.375 input and $1.875 output, half the standard rate. Their 2027 prices match today's standard prices, so batch workloads absorb the increase.

Do thinking tokens count as output? Yes. Google prices output "including thinking tokens," and says the model may spend more of them on complex tasks at higher effort levels. Lower the effort level to cut token overhead.

What is Gemini 3.8 Flash Cyber and what does it cost? A cybersecurity variant restricted to trusted defenders through Google's Fairwind Program. No public pricing, since access is by application rather than by API key.

What is the context window? 1,048,576 tokens, per OpenRouter's model listing for both the standard and batch endpoints.

Conclusion

Gemini 3.8 Flash at $0.75 and $3.75 is a good rate for a model this capable, and it is worth switching to from 3.6 or 3.7 because those cost the same. Take the free upgrade.

The number to write down is 1 January 2027, when every line of that table doubles. If a meaningful share of your Flash traffic can run asynchronously, batch it before then and the increase never reaches you.

Sources

Tools In This Article

Browser-based, no sign-up. Try them while the topic is fresh.

More From ToolMintX

Other Blog Posts