Mistral released a public preview of Mistral Large 4 on 6 October 2026. It is a mixture-of-experts model with about 1.05 trillion total parameters and 52 billion active per token, a 1M-token context window and image input. The API model name is mistral-large-4, and Mistral says the open weights come out at the end of October.
Mistral's own pages show two prices for it. The launch post lists $1.36 per million input tokens and $4.18 per million output. The model page in the docs lists $0.68 input, $0.07 cached input and $2.09 output, with the $1.36 / $0.14 / $4.18 figures shown as the original price. So the preview runs at exactly half the list rate, and neither page says when that ends.
This post costs a realistic coding-agent session at both rates, puts it next to DeepSeek V4 Pro, GLM-5.3, GPT-6.1 Sol and Claude Sonnet 5.5, and covers what changes when the weights land.
Table of Contents
- The two prices on Mistral's pages
- One agent session, seven price points
- What the benchmarks say, and who ran them
- Can you run it yourself when the weights drop?
- Three things to check before you switch
- FAQ
- Conclusion
- Sources
The two prices on Mistral's pages
Per million tokens:
| Rate | Input | Cached input | Output |
|---|---|---|---|
| Preview (docs model page) | $0.68 | $0.07 | $2.09 |
| List ("original price" on the same page, and the launch post) | $1.36 | $0.14 | $4.18 |
Every line is cut by the same 50%, cache reads included. That matters for agents, because cache reads are most of what an agent pays for. It also means the rate you budget at today may double later. Mistral calls the model a public preview and says the reinforcement learning run "is still in flight", so the model and the price could both change before general availability. The docs page also carries the caveat that "the price may change depending on the features used".
For comparison, Mistral's previous flagship, Mistral Large 3, is listed at $0.50 input and $1.50 output. Mistral Medium 3.5 is $1.50 / $7.50. Large 4 at preview pricing costs a little more than Large 3 per token, and at list pricing it costs a bit under Medium 3.5 on input and well under it on output.
One agent session, seven price points
I used the same session as our Sonnet 5.5 vs GPT-6.1 Sol comparison so the numbers line up:
- 60 turns
- a 120,000-token prefix (system prompt, tool schemas, files read), written to the cache on turn 1 and re-read on the other 59 turns
- 3,000 new input tokens and 1,500 output tokens per turn
That is 7.38 million input tokens, 95.9% of them cache reads, plus 90,000 output tokens. Mistral and Z.ai don't list a separate cache-write price, so the first write is billed at the input rate. OpenAI and Anthropic charge $2.50 for that 5-minute cache write.
| Model | Cache write | Cache reads | New input | Output | Session | Per month* |
|---|---|---|---|---|---|---|
| DeepSeek V4 Pro (off-peak) | $0.08 | $0.16 | $0.12 | $0.18 | $0.53 | $234 |
| Mistral Large 4 (preview) | $0.08 | $0.50 | $0.12 | $0.19 | $0.89 | $391 |
| DeepSeek V4 Pro (peak) | $0.16 | $0.31 | $0.24 | $0.36 | $1.06 | $468 |
| Mistral Large 4 (list) | $0.16 | $0.99 | $0.24 | $0.38 | $1.78 | $781 |
| GPT-6.1 Sol | $0.30 | $0.71 | $0.36 | $0.90 | $2.27 | $998 |
| GLM-5.3 | $0.17 | $1.84 | $0.25 | $0.40 | $2.66 | $1,169 |
| Claude Sonnet 5.5 | $0.30 | $1.42 | $0.36 | $0.90 | $2.98 | $1,309 |
*20 sessions a day for 22 working days (440 sessions). Totals are computed from unrounded rates, so a row can differ from its rounded columns by a cent.
At the preview rate, Large 4 costs less than a third of Sonnet 5.5 for this session and about 39% of GPT-6.1 Sol. At list price the gap narrows, but it is still the cheaper option against both. GLM-5.3 is the open-weight model Mistral compares against most often in its post, and it costs more than Large 4 at either rate because its cached input is $0.26.
DeepSeek V4 Pro stays cheapest. Its cache hit is $0.022 off-peak, about a third of Large 4's preview rate. Peak hours (01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays) double every DeepSeek rate. That covers 6:30 to 9:30 AM and 11:30 AM to 3:30 PM in India, which is a large part of a working day here.
Caching matters more than the model. Without it, the same session costs $5.21 on Large 4 at preview rates, nearly six times the cached cost. If your agent framework rewrites the system prompt every turn or a proxy strips the cache, fix that before you compare models.
You can plug in your own token counts in the API Cost Calculator, which now lists Mistral Large 4 at the preview rate with the list rate in its note.
What the benchmarks say, and who ran them
All of these numbers come from Mistral's launch post. I haven't seen independent reproductions yet.
- Coding: 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4. Mistral puts its combined Coding Agent Index at 49.8%, ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max.
- Blind human eval (run with Surge AI): second of five on coding quality at 3.74 out of 5, ahead of GLM-5.3 (3.60) and Kimi K3 (3.59), behind Claude Opus 5 (4.22).
- Agents: 59.9% on AutomationBench, 657 business workflows across apps like Gmail, Sheets, Slack and Salesforce.
- Security: 93% on Cybench, and 82% on a reproduce-then-patch vulnerability test where Mistral says Claude Opus 5.5 and GPT-6 Astra score near zero because they refuse the task.
Compare the Terminal-Bench number with what Anthropic reported for Sonnet 5.5: 70.6% on Terminal-Bench 4.0. Large 4 is a cheaper model, but it is a long way behind on long terminal tasks. If your agent spends most of its time running shell commands and reading their output, the lower bill may get eaten by extra turns and retries. Measure turns per finished task on your own repo, then cost those turns.
Can you run it yourself when the weights drop?
Not on a workstation. The docs model page gives 1.05T total parameters and 52B active, plus a 1.6B vision encoder. The launch post rounds that to 1 trillion and 49 billion active, so expect the final card to settle the exact figure. The GPU memory column on the docs page currently reads N/A, and the license column reads "Coming soon".
Weights alone, before any KV cache:
| Precision | Bytes per parameter | Weights for 1.05T params |
|---|---|---|
| BF16 | 2 | about 2,100 GB |
| FP8 | 1 | about 1,050 GB |
| 4-bit | 0.5 | about 525 GB |
A mixture-of-experts model has to keep all its experts in memory even though each token only uses about 5% of them. Active parameters set the speed. Total parameters set the memory you need. Even the 4-bit version needs something like eight 80 GB GPUs once you leave room for context. The AI VRAM Calculator will give you the full figure with KV cache once Mistral publishes the config.
So for most teams the open weights are about control, not about saving money. They let a bank or a government run the model on its own servers under its own rules, which is the pitch Mistral makes in the post. If you just want a cheap agent model, the API is the realistic route.
Three things to check before you switch
1. Budget at the list rate. The half-price preview has no end date on either page. If your plan only works at $0.89 a session, it might stop working without much warning. At $1.78 it still beats GPT-6.1 Sol and Sonnet 5.5 on this workload.
2. Reasoning effort defaults to high. The docs code sample passes reasoning_effort="high", and the model supports high or none. Reasoning tokens bill as output. For simple tool-routing turns, try none and compare quality, since output is the most expensive line per token.
3. It talks the Mistral API, not the OpenAI one. The model works through /v1/chat/completions and /v1/conversations with function calling, structured outputs and batching. Batch jobs are listed at a 50% discount, which helps overnight evals. If your agent hardcodes OpenAI's Responses API, you need a different client or an adapter. Validate a few tool-call payloads with the JSON Formatter before trusting a long run.
FAQ
How much does Mistral Large 4 cost? During the public preview, $0.68 per million input tokens, $0.07 cached input and $2.09 output. The list price shown on the same docs page and in the launch post is $1.36, $0.14 and $4.18.
When does the Mistral Large 4 preview price end? Mistral hasn't said. The docs show the higher figures as the original price and give no end date, so plan for the list rate.
Is Mistral Large 4 open source? Mistral calls it open-weight and says the weights will be released by the end of October 2026. The license is listed as "Coming soon", so check it before you plan commercial self-hosting.
How much VRAM does Mistral Large 4 need? About 2,100 GB for the weights at BF16 and about 525 GB at 4-bit, based on 1.05T total parameters. That is a multi-GPU server, not a desktop.
Is Mistral Large 4 cheaper than Claude Sonnet 5.5 for coding agents? Yes on price. A 60-turn cached session costs $0.89 at preview rates and $1.78 at list, against $2.98 on Sonnet 5.5. Sonnet 5.5 scores much higher on Terminal-Bench 4, so it can finish in fewer turns.
Conclusion
Mistral Large 4 is cheap to try right now. A typical cached agent session costs $0.89, under a third of Sonnet 5.5. The catch is the preview label: the price is half of a list rate that has no published switch date, the model is still being trained, and the coding benchmarks Mistral picked put it behind the closed models on terminal work. Budget at $1.78, test turns per task on your own code, and treat the open weights as a self-hosting option for large organisations. Run your own numbers in the API Cost Calculator.
Sources
- Mistral AI, "Introducing Mistral Large 4" (6 October 2026): preview status, weights release timing, parameter counts, benchmarks, human eval, training details and the $1.36 / $4.18 list price.
- Mistral Docs, Mistral Large 4 model page (v26.10): 1.05T total / 52B active / 1.6B vision encoder, 1M context, preview price $0.68 / $0.07 / $2.09 with original price $1.36 / $0.14 / $4.18, supported endpoints and the
reasoning_effortusage sample. - Mistral Docs, Mistral Large 3 and Mistral Medium 3.5 model pages: $0.50 / $1.50 and $1.50 / $7.50.
- DeepSeek API Docs, Models & Pricing: DeepSeek V4 Pro peak and off-peak rates and peak hours.
- Z.ai Developer Docs, Pricing: GLM-5.3 rates.
- OpenAI and Anthropic pricing pages, as cited in our 3 October comparison: GPT-6.1 Sol and Claude Sonnet 5.5 rates and Sonnet 5.5's Terminal-Bench 4.0 score.
- Session and memory figures are calculated by ToolMintX from those rates and parameter counts.
