IT Tool

AI API Cost Calculator

Estimate and compare API costs for popular AI models across providers — text, image, video, and embeddings — by input/output token pricing.

Instant 100% Client-Side No Login
PROCESSINGLOCAL
LIMITNONE
PRIVACYBROWSER-ONLY
Show costs in
Prices gathered September 29, 2026 and are per 1M tokens unless noted. API pricing changes often — always confirm on the provider’s official page before relying on a figure.

Cost / request

$0.0100

Total (1,000 requests)

$10.00

Selected model

GPT-5.4

All text models — cost for your inputs

· cheapest first
ModelProviderCost / requestTotal × 1,000
Mistral SmallMistral$0.000250$0.2500
GLM-4.7-FlashXZ.ai (GLM)$0.000270$0.2700
GPT-4.1 nanoOpenAI$0.000300$0.3000
Gemini 2.5 Flash-LiteGoogle (Gemini)$0.000300$0.3000
GPT-6 LunaReleased Sep 22, 2026 as OpenAI's most efficient tier for high-volume tasks. 1.05M context. Cache writes $0.125. Same 2x input / 1.5x output penalty above 272K input tokens.OpenAI$0.000350$0.3500
GLM-5.3-FlashList price, now in effect. The 50% launch promo ($0.075 in / $0.25 out / $0.015 cached) ended at 24:00 on Sep 9, 2026 (UTC+8). 320B-A18B MIT-licensed open weights, 1M context.Z.ai (GLM)$0.000400$0.4000
GPT-4o miniOpenAI$0.000450$0.4500
Grok 4 FastxAI (Grok)$0.000450$0.4500
DeepSeek V4 Flash (chat/reasoner)Off-peak rate (since Aug 16, 2026). Peak: $0.44 in / $1.32 out per 1M.DeepSeek$0.000550$0.5500
GPT-5.6 LunaCheapest GPT-5.6 tier (nano).OpenAI$0.000800$0.8000
GPT-5.4 nanoOpenAI$0.000825$0.8250
GPT-4.1 miniOpenAI$0.001200$1.20
Gemini 2.5 FlashGoogle (Gemini)$0.001550$1.55
DeepSeek V4 ProOff-peak rate (since Aug 16, 2026). Peak: $1.32 in / $3.96 out per 1M.DeepSeek$0.001650$1.65
Gemini 3 FlashGoogle (Gemini)$0.002000$2.00
Grok Build 0.1256K context. Above 200K prompt tokens: $2 in / $0.40 cached / $4 out.xAI (Grok)$0.002000$2.00
Grok 4.31M context. Above 200K prompt tokens: $2.50 in / $0.40 cached / $5 out. 20% batch discount available.xAI (Grok)$0.002500$2.50
Gemini 3.8 FlashLaunched Sep 2, 2026. Introductory rate through Dec 31, 2026; $1.50/$7.50 from Jan 1, 2027. Batch and Flex are half these rates. Output includes thinking tokens.Google (Gemini)$0.002625$2.62
Gemini 3.7 FlashIntroductory rate through Dec 31, 2026; $1.50/$7.50 from Jan 1, 2027. Output includes thinking tokens.Google (Gemini)$0.002625$2.62
Gemini 3.6 FlashLaunched Jul 2026. Same introductory rate as 3.8 Flash through Dec 31, 2026; $1.50/$7.50 from Jan 1, 2027. Output includes thinking tokens.Google (Gemini)$0.002625$2.62
GPT-5.4 miniOpenAI$0.003000$3.00
Claude Haiku 4.5Anthropic (Claude)$0.003500$3.50
GLM-5.3Z.ai (GLM)$0.003600$3.60
GLM-5.2Z.ai (GLM)$0.003600$3.60
Mistral LargeMistral$0.005000$5.00
Grok 4.7Released Sep 21, 2026 at the same price as Grok 4.6. 500K context. Above 200K prompt tokens the whole request bills at $4 in / $1 cached / $12 out.xAI (Grok)$0.005000$5.00
Grok 4.6500K context, configurable reasoning. Above 200K prompt tokens: $4 in / $1 cached / $12 out.xAI (Grok)$0.005000$5.00
Grok 4.5Above 200K prompt tokens: $4 in / $0.60 cached / $12 out.xAI (Grok)$0.005000$5.00
GPT-4.1OpenAI$0.006000$6.00
Gemini 3.5 FlashGoogle (Gemini)$0.006000$6.00
Gemini 2.5 ProHigher rate ($2.50/$15) above 200k input tokens.Google (Gemini)$0.006250$6.25
GPT-6 SolReleased Sep 22, 2026 for complex coding and agentic workflows. 1.05M context, 128K max output. Cache writes $2.50. Prompts above 272K input tokens are billed at 2x input and cache rates and 1.5x output for the whole request.OpenAI$0.007000$7.00
Claude Sonnet 5.5Listed on the official pricing table in Sep 2026 at the same base rate as Sonnet 5. Cache writes $2.50 (5 min) / $4 (1 hour). Batch API $1 in / $5 out.Anthropic (Claude)$0.007000$7.00
Claude Sonnet 5Still $2/$10 on the official pricing table as of Sep 1, 2026, past the original Aug 31 intro cutoff.Anthropic (Claude)$0.007000$7.00
GPT-4oOpenAI$0.007500$7.50
Cohere Command ACohere$0.007500$7.50
GPT-5.6 TerraBalances intelligence and cost (mini tier).OpenAI$0.008000$8.00
Gemini 3.1 ProHigher rate ($4/$18) above 200k input tokens.Google (Gemini)$0.008000$8.00
GPT-5.4OpenAI$0.0100$10.00
Claude Sonnet 4.6Anthropic (Claude)$0.0105$10.50
Grok 4xAI (Grok)$0.0105$10.50
GPT-5.6 SolTop GPT-5.6 tier, superseded by GPT-6 Sol on Sep 22, 2026 at half these rates. Promotional pricing through Nov 21, 2026 (dropped 20%+ on Aug 21, 2026).OpenAI$0.0140$14.00
Claude Opus 5.5Released Sep 22, 2026; first of the Claude 5.5 family and the current flagship. 1M context, 128K max output. Cache writes $5 (5 min) / $8 (1 hour). Fast mode $8/$40 for up to 2.5x speed. Anthropic's "40% cheaper than Opus 5" figure combines this 20% price cut with fewer tokens used per task, so per-token savings alone are 20%.Anthropic (Claude)$0.0140$14.00
Claude Opus 5Superseded by Opus 5.5 on Sep 22, 2026. 1M context at no premium. Cache writes $6.25. Fast Mode $10/$50.Anthropic (Claude)$0.0175$17.50
Claude Opus 4.8Anthropic (Claude)$0.0175$17.50
GPT-5.5OpenAI$0.0200$20.00
GPT-6 AstraReleased Sep 3, 2026. Standard rate; Fast mode costs 2x for up to 2x speed. Cache writes $12.50. Long-context requests bill at $20 in / $2 cached / $75 out (2x input, 1.5x output).OpenAI$0.0350$35.00
Claude Fable 5.1Sep 1, 2026. Same base price as Fable 5; cache reads cut from $1 to $0.25/MTok.Anthropic (Claude)$0.0350$35.00
Claude Mythos 5.1Same capabilities and price as Fable 5.1; invitation only (Project Glasswing).Anthropic (Claude)$0.0350$35.00
Claude Fable 5Anthropic (Claude)$0.0350$35.00
GPT-5.6 CyberHighest-capability GPT-5.6 tier.OpenAI$0.0500$50.00
Claude Opus 4.1Deprecated.Anthropic (Claude)$0.0525$52.50
GPT-5.5 ProOpenAI$0.1200$120.00

The headline price is per million tokens, and that is what fools people

Almost every language model API bills by the token, not by the request. A token is a chunk of text — roughly four characters, or about three-quarters of an English word — and both the prompt you send and the completion the model returns are counted. Prices are quoted per million tokens, which makes the numbers look trivially small: $2.50 per million sounds like nothing until you notice a single retrieval-augmented chat turn can carry 20,000 tokens of context, and that you are serving it tens of thousands of times a day. This calculator exists to turn that per-million rate into the two figures that actually drive a budget — cost per request, and cost across your real monthly volume — and to line every model up in one sorted table so the cheap option for your specific token mix is obvious rather than buried in seven different pricing pages.

Output is the line item that bites

The single most important thing to understand is that input and output tokens are priced separately, and output almost always costs several times more because generating text is far more compute-intensive than reading it. GPT-5.4 is $2.50 per million in and $15 out; Claude Opus 4.8 is $5 in and $25 out; even DeepSeek V4 Flash keeps the 2× gap at $0.14 and $0.28. The practical consequence is that a model with a cheap input rate can still lose on total cost if your workload is generation-heavy. A summariser that reads 10,000 tokens and writes 200 is dominated by input; a code generator that reads 500 and writes 4,000 is dominated by output. Look at the column that matches your traffic, not the one in the headline.

Prompt caching is the biggest discount most teams miss

When you send the same large context on every call — a long system prompt, a fixed knowledge-base document, or a growing conversation history — providers like OpenAI and Anthropic will bill the repeated portion at roughly a tenth of the normal input rate on a cache hit. GPT-5.4 drops from $2.50 to $0.25 per million on cached input; Claude Opus 4.8 from $5 to $0.50. For an agent that resends a 15,000-token instruction block on every one of thousands of turns, that alone can cut the input bill by the better part of an order of magnitude. The text tab has a dedicated cached-tokens field so you can see the split between fresh and cached input reflected in the estimate rather than guessing at it.

Image, video, and embeddings are priced on completely different units

Not every modality is a token. Image APIs mostly charge a flat rate per generated image — FLUX.1 Schnell at about $0.003, Gemini 2.5 Flash Image at $0.039, Ideogram v3 at $0.09 — a 30× spread that matters enormously at volume. Video is billed per second of output, from Veo 3.1 Lite at $0.05 to Veo 3.1 Standard at $0.40, so a single 60-second clip ranges from $3 to $24 depending only on the model. A few image models, such as OpenAI's GPT Image 2, are priced by image tokens rather than per image; because this tool has no image-token input to price them honestly, they are deliberately left out of the per-image comparison rather than shown as a misleading zero. Embeddings are the cheapest line on any AI bill — OpenAI's text-embedding-3-small is $0.02 per million tokens — which is why embedding a whole corpus for search usually costs less than a single day of chat traffic.

How to actually cut the bill, and where these numbers stop

The cheapest model is rarely the right default and the flagship is rarely necessary. The highest-leverage move is to route by difficulty: send classification, extraction, and simple summarisation to a mini or nano tier, Gemini Flash-Lite, or DeepSeek, and reserve GPT-5.5, Claude Opus, or Gemini Pro for genuine reasoning and high-stakes output. Because the gap between tiers is often 10× to 50×, shifting routine traffic downward usually saves more than any other single change. Estimate before you build: put your realistic average token counts and monthly volume into the calculator, compare the top few candidates, and you will frequently find two models of similar quality that differ several-fold in cost. What this tool deliberately does not model is everything beyond the base per-token rate — batch tiers (commonly around 50% off), regional and data-residency endpoints, fine-tuning surcharges, and server-side tools like web search all change the final invoice, and every provider's live pricing page is linked in the tool so you can confirm the current figure before you commit budget. The prices here were gathered in September 29, 2026 and are a planning aid, not a contract.

How to Use

1

Pick a modality: text/chat, image, video, or embeddings.

2

Choose a model, then enter your token counts (or images/seconds) and number of requests.

3

See cost per request and total, plus a comparison table of every model sorted cheapest-first.

4

Cross-check the number on the provider’s official pricing page before you rely on it.

Features

Text, image, video, and embedding pricing in one place
Compare every model side by side, sorted cheapest-first
Handles input, output, and cached (prompt-cache) token pricing
Covers OpenAI, Claude, Gemini, DeepSeek, Mistral, Grok, and Cohere
100% client-side — nothing is sent to a server

Common Questions

About AI API Cost Calculator

Estimate and compare the cost of AI API calls across every popular provider and model. Enter input and output token counts to see the cost per request and monthly total for OpenAI GPT, Anthropic Claude, Google Gemini, DeepSeek, Mistral, xAI Grok, and Cohere. Includes prompt-cache pricing (cached input billed at a fraction of the normal rate), image generation cost per image, video generation cost per second, and embedding cost per token, with a comparison table sorted cheapest-first and optional USD/INR display. All calculations run in your browser.

Also known as: llm api cost, token cost calculator, openai pricing calculator, ai api pricing, gpt cost per token, claude pricing calculator, gemini api cost, deepseek pricing, prompt caching cost, embedding cost calculator, image generation api cost, video generation api cost, compare llm prices, ai api price comparison.

Processing Note

AI API Cost Calculator runs in your browser, so the input you enter is processed locally on this page and is not uploaded to a ToolMintX account.

Tool Limits

IT tools provide quick diagnostics and transformations. They cannot see every private network, deployment setting, proxy, firewall, or production edge case.

Explore More