Why 39% fewer tokens cost us 149.6% more, and what we changed

Saving6 min read

If you run a coding agent all day, you have likely tried to cut its token count and watched the bill stay the same, or go up. We did too. Prompt caching already makes most of each request cheap. The next saving comes from keeping the cached part of the request still, trimming what sits in it, and pricing each request the way the provider prices it. Here is what Context Mode does for each, what we measured, and what we have not measured yet.

The short version

The problem: the bill does not follow the token count

A coding agent sends the whole conversation on every call. Prompt caching makes most of that cheap. On Claude Opus 5.5 the list price per million tokens is:

Token typePrice per million tokens
Fresh input$4
1-hour cache write$8
Cache read$0.20

A read costs a twentieth of a fresh token. A write costs forty times a read. That ratio decides whether a change saves money: if it makes the provider write the prefix again, it has to remove a lot of read tokens to pay that write back. In our August 2026 benchmark grid on Claude Haiku 4.5 (write $2.00, read $0.10, so 20:1), cache writes were 45 to 94% of each cell's cost.

On 2026-10-01 we tested tool search: send none of the 60 MCP tool definitions up front, and let the model search for the ones it needs. Prompt tokens per session fell 39%, from 176,196 to 107,495. Every task still passed, 12 of 12 in both arms.

Cost rose 149.6% over 12 paired runs on Opus 5.5.

Tool search off$0.1043
Tool search on, 39% fewer tokens$0.2107
Mean cost per run, 12 paired runs on Opus 5.5, 60 MCP tools. Cost rose 149.6% (95% interval +104.1 to +195.1%).

The 60 definitions had been cheap cache reads. Taking them out caused new cache writes and uncached input. We did not ship it. Everything below cuts tokens without breaking the cache.

Fix 1: trim tool schemas, keep every tool

Tool schemas sit at the head of every request. In one recording, Claude Code sent a 29-tool manifest of 103,253 bytes.

Context Mode does not remove tools. Each tool of 1,500 bytes or more keeps its name, its top-level parameters and its required list. Its description is cut at a sentence end at or before 420 characters, with a note that tells the model to call context-mode-tool-schema for the rest. The slim form depends only on the tool list, so it is the same bytes on every turn.

We measured it with a recording proxy in front of a real claude -p on a Haiku model, 7 turns:

Full manifestSlim manifest
Manifest size103,253 B27,436 B
Prompt tokens, warm turn40,68220,748
Cost per warm turn$0.004596$0.002640 (-42.5%)
The turn the manifest changes$0.040917

Look at the last row. The turn where the manifest changes costs 8.9 times a warm turn, because the provider writes 20,143 tokens to cache again. That one turn takes about 20 warm turns to pay back. A trim that changed its output from turn to turn would never pay back. That is why the trim must give the same bytes every time.

The same study tried the other way first: removing 26 tools instead of trimming them made that turn 4.7 times more expensive.

Now on Codex too

Codex 0.156.1 declares its tools in a different shape, and until now Context Mode did not trim them. It does now. The Codex trim cuts descriptions only:

On one captured gpt-6-astra request, with two live requests, input fell from 15,347 to 14,083 tokens: 1,264 fewer per request. In a live check, real codex exec turns that each had to call a tool passed 5 of 5 on gpt-6-astra and 3 of 3 on gpt-5.5.

Keep the size in proportion. After the first request, those 1,264 tokens are cache reads, so the dollar effect per turn is small. Context Mode also adds its own tokens to a Codex request: on 2026-09-25 we measured 1,799 more input tokens on one forwarded request, with memory and skills off. That was a different capture, so do not subtract one number from the other.

Fix 2: hold the prefix still

In our own data, this mattered most.

In the 7 days to 2026-10-01, one heavy session spent $657.22 on its main agent. Of that, $331.30 (50.4%) was cold cache writes: requests that wrote their prompt to cache again. Almost every request used the 1-hour cache, yet 54 of the 65 cold writes came less than an hour after the previous request. The cache did not expire. Something changed the prefix.

So we found the cause of every cold write that came before the cache could expire, across three tenants, from 2026-09-22 to 2026-10-02: 198 cold writes. These are the causes we can name, plus the unknown ones:

CauseCold writes
Context Mode moved the point where it cuts old messages90
The client changed model or cache TTL (for example /model)18
Unknown, all before the fix44

The largest cause was ours. When a long thread nears the context window, Context Mode cuts old messages to keep it under the limit, and that cut point kept moving. The fix holds the cut in place and moves it only in planned steps:

The fix has been on by default since 2026-10-03. The rule we took from it: every change Context Mode makes to a request is a pure function of its input, made once, and the same bytes after.

Fix 3: count the price tier (Claude Haiku 5.5)

Claude Haiku 5.5 prices a whole request at one of two tiers. At or below 100,000 prompt tokens: $0.10 input, $0.50 output, $0.01 cache read, $0.20 1-hour cache write. Above that, every rate is 5 times higher. The prompt length counts all input tokens, cache reads and writes included.

So a request of 100,001 tokens pays 5 times for every token, the first 100,000 too. If Context Saving shortens a request from above the line to below it, the lower tier is a saving on top of the tokens removed.

What shipped on 2026-10-09: the ledger prices each call at its own tier and books the tier drop as its own row, "Lower price tier". In a unit test, a request folded from 150,003 to 60,003 tokens is priced at the low tier, and the tier drop alone is 4 times the bill. That is a test case, not a live measurement.

What did not ship: Context Mode does not aim at the line. It does not fold more because a request is near 100,000 tokens. We have no Haiku 5.5 runs through it yet.

Fix 4: count Codex cost per request

You cannot cut what you do not count. Codex cost is now measured per request:

What is not measured yet

How to try it

Context Mode sits between your coding agent and the model. It works with Claude Code and Codex today; OpenCode and Pi are coming soon. Your own Anthropic or OpenAI account pays the model.