Editorial

Sonnet 5.5 vs GPT-6.1 Sol: Cache Reads Decide the Agent Bill

Same $2/$10 rates, different agent bills. A 60-turn coding session costs $2.98 on Sonnet 5.5 and $2.27 on GPT-6.1 Sol. Why.

JJyoti Ranjan SwainUpdated
Claude Sonnet 5.5 and GPT-6.1 Sol both at $2 input and $10 output, with cache reads at $0.20 and $0.10

Two mid-tier models shipped a day apart: Claude Sonnet 5.5 on 28 September and GPT-6.1 Sol on 29 September. Both list at $2 per million input tokens and $10 per million output. On the pricing table they look like the same model at the same price.

They don't cost the same once you run a coding agent on them. In an agent session almost all of the input is the same long prefix (system prompt, tool definitions, the files already read), sent again on every turn. That repeated prefix bills at the cache-read rate, and the two models charge different cache-read rates: GPT-6.1 Sol's is half of Sonnet 5.5's. This post works through one agent session on both, shows where the money goes, and covers the two workflow catches that the headline price leaves out.

Table of Contents

The rates, side by side

Per million tokens, from each provider's official pricing page:

ModelInputCache readCache writeOutput
Claude Sonnet 5.5$2.00$0.20$2.50 (5 min), $4 (1 hour)$10.00
GPT-6.1 Sol$2.00$0.10$2.50$10.00
GPT-6 Sol$2.00$0.20$2.50$10.00
Claude Opus 5.5$4.00$0.20$5.00 (5 min), $8 (1 hour)$20.00

The one number that differs between the two new models is the cache read: $0.10 against $0.20. GPT-6.1 Sol also halved it compared with GPT-6 Sol, which launched a week earlier at $0.20.

Both have a 1M-token context window and 128K max output. Above that the pricing splits. Anthropic bills its full 1M window at the standard rate, so a 900K-token request costs the same per token as a 9K one. OpenAI charges more once a prompt goes past 272K input tokens: GPT-6.1 Sol rises to $4 input, $0.20 cache read, $5 cache write and $15 output for that request.

Why cache reads decide an agent bill

A coding agent like Claude Code, Codex or Hermes doesn't send one prompt. It sends the whole conversation again on every turn: the system prompt, every tool schema, every file it has opened, every command output so far. Turn 40 carries everything from turns 1 to 39.

Prompt caching means you pay full input price on that prefix only once, when it is written to the cache. After that, each turn re-reads it at the cache-read rate. The only tokens billed at full input price each turn are the new ones: your next message and the latest tool result.

So the share of input that bills at the cache-read rate is very high. In the session below it is 95.9% of all input tokens. If one model charges half as much for that 95.9%, its total bill drops sharply, even though the input and output rates printed on the table are identical.

Cost of one 60-turn coding session: Opus 5.5 $4.54, Sonnet 5.5 $2.98, GPT-6 Sol $2.98, GPT-6.1 Sol $2.27

One coding session, four models

Take a realistic session:

  • 60 turns
  • a 120,000-token prefix (system prompt, tools and files read), written to the cache on turn 1 and re-read on the other 59
  • 3,000 new input tokens per turn
  • 1,500 output tokens per turn

That is 7.38 million input tokens in total. Here is what each model charges:

ModelCache writeCache readsNew inputOutputTotal
GPT-6.1 Sol$0.30$0.71$0.36$0.90$2.27
Claude Sonnet 5.5$0.30$1.42$0.36$0.90$2.98
GPT-6 Sol$0.30$1.42$0.36$0.90$2.98
Claude Opus 5.5$0.60$1.42$0.72$1.80$4.54

Every column is the same for Sonnet 5.5 and GPT-6.1 Sol except cache reads. That single column makes GPT-6.1 Sol $0.71 cheaper per session, 23.8% less. Opus 5.5 costs 1.52 times as much as Sonnet 5.5.

Now remove caching. The same session with every turn paying full input price on the prefix comes to $15.66 on Sonnet 5.5 or GPT-6.1 Sol and $31.32 on Opus 5.5. That is five to seven times the cached cost. If your agent framework, a proxy, or a changing system prompt breaks the cache, the choice of model matters less than that bug.

At 20 sessions a day for 22 working days, the monthly gap adds up: about $998 on GPT-6.1 Sol, $1,309 on Sonnet 5.5 and $1,996 on Opus 5.5.

You can run your own token counts through the API Cost Calculator, which now includes GPT-6.1 Sol with its cache rates.

What Sonnet 5.5 changed besides price

The per-token numbers didn't move from Sonnet 5, so Anthropic's "up to 30% less" claim is about using fewer tokens per task, not a lower rate. Its announcement says Sonnet 5.5 batches more tool calls together and finishes tasks in fewer steps. Fewer turns means fewer cache reads, which is the line that dominates the table above. If you get the 30% saving in practice, the Sonnet 5.5 session drops to roughly $2.08. That is below GPT-6.1 Sol's $2.27, but only if your workload behaves the way Anthropic's tests did. Measure it on your own tasks.

Other points from Anthropic's announcement and docs:

  • Terminal-Bench 4.0: 70.6%, against Sonnet 5's 10.3% and Opus 5.5's 66.4% in Anthropic's own table. On CursorBench 4.0 it scores 55.5% against Opus 5.5's 57.8%.
  • Speed: outputs generate more than 30% faster than Sonnet 5.
  • Default effort is high for Sonnet 5.5, while Opus 5.5 defaults to medium. If you are comparing the two on cost, set the effort explicitly so you compare like for like.
  • Haiku 5.5 is announced for "the coming weeks", so there will be a cheaper tier in the same family.

Anthropic is also clear that Opus 5.5 "remains clearly stronger at complex, open-ended work requiring sustained judgment". The benchmark lead on Terminal-Bench doesn't make Sonnet 5.5 a replacement for Opus 5.5 on hard tasks.

Two catches with GPT-6.1 Sol

1. Tool calling needs the Responses API. OpenAI's model page says Chat Completions is supported without tool calling. A coding agent is mostly tool calls, so if your framework talks to OpenAI through /v1/chat/completions, GPT-6.1 Sol will answer but won't run tools. Check that your agent uses /v1/responses before you switch.

2. The 272K line doubles the input price. Agents that load a whole repository into context can cross 272K input tokens. A single 300K-token request with 4K of output costs about $0.64 at the standard rate, but GPT-6.1 Sol bills it at $1.26 because the long-context rates apply. Sonnet 5.5 has no such line. If your sessions regularly grow past 272K, the cache-read saving can disappear.

GPT-6.1 Sol also doesn't support the none or minimal reasoning efforts. The lowest setting is low.

Which one to put in your agent loop

  • Long sessions that stay under 272K and already use the Responses API: GPT-6.1 Sol. The cheaper cache read is a real, structural saving of about a quarter on a typical session.
  • Very long contexts, or an agent built on Chat Completions: Sonnet 5.5. No long-context penalty and no API change needed.
  • Hard, open-ended work where a wrong step is expensive: Opus 5.5. It costs about 1.5 times as much per session, which is cheap compared with redoing a broken change.
  • Any of them: make sure caching actually works. A broken cache costs five to seven times more than the model choice saves.

A practical setup is to route by task: Sonnet 5.5 or GPT-6.1 Sol for routine fixes and reviews, Opus 5.5 for the hard ones. That is the workflow Anthropic itself describes for the 5.5 family.

FAQ

Are Claude Sonnet 5.5 and GPT-6.1 Sol the same price? They have the same $2 input and $10 output rates. GPT-6.1 Sol's cache read is $0.10 per million tokens against Sonnet 5.5's $0.20, so it is cheaper for agent workloads that re-read a long cached prefix.

How much cheaper is GPT-6.1 Sol for a coding agent? In a 60-turn session with a 120K-token cached prefix it costs $2.27 against Sonnet 5.5's $2.98, which is 23.8% less. The gap depends on how much of your input is cached.

Did Sonnet 5.5 get cheaper than Sonnet 5? Not per token. It is priced the same as Sonnet 5. Anthropic says it uses fewer tokens per task, up to 30% less in its testing.

Does GPT-6.1 Sol support tool calling in Chat Completions? No. OpenAI says Chat Completions is supported without tool calling. Use the Responses API for agents.

Is there a long-context surcharge on Claude Sonnet 5.5? No. Anthropic bills the full 1M-token window at standard rates. GPT-6.1 Sol charges higher rates above 272K input tokens.

Conclusion

The headline rates for Claude Sonnet 5.5 and GPT-6.1 Sol are identical, and for chat-style use they cost about the same. For coding agents the bill is mostly cache reads, and GPT-6.1 Sol charges half as much for them, which makes a typical session about a quarter cheaper. Sonnet 5.5 can win that back if it really uses fewer turns, and it has no long-context surcharge or API restriction. Run your own numbers in the API Cost Calculator, and if you are also sizing a local model for the same work, the AI VRAM Calculator covers that side.

Sources

  • Anthropic, "Introducing Claude Sonnet 5.5" (28 September 2026): pricing, Terminal-Bench 4.0, CursorBench 4.0 and speed figures, and the Haiku 5.5 note.
  • Claude Platform Docs, Pricing: Sonnet 5.5 and Opus 5.5 rates, cache write and read rates, and long-context pricing at standard rates for Claude 4.6 and later.
  • Claude Platform Docs, Models overview: context window, max output and default effort for Sonnet 5.5 and Opus 5.5.
  • OpenAI API changelog (29 September 2026): GPT-6.1 Sol release and standard pricing.
  • OpenAI API, GPT-6.1 Sol model page: context window, max output, reasoning efforts, and Chat Completions without tool calling.
  • OpenAI API pricing page: GPT-6.1 Sol and GPT-6 Sol short-context and long-context rates.
  • Session costs are calculated by ToolMintX from those rates using the session described above (60 turns, 120K cached prefix, 3K new input and 1.5K output per turn).

Tools In This Article

Browser-based, no sign-up. Try them while the topic is fresh.

More From ToolMintX

Other Blog Posts