# Why 39% fewer tokens cost us 149.6% more, and what we changed

October 9, 2026 Saving 6 min read

If you run a coding agent all day, you have likely tried to cut its token count and watched the bill stay the same, or go up. We did too. Prompt caching already makes most of each request cheap. The next saving comes from keeping the cached part of the request still, trimming what sits in it, and pricing each request the way the provider prices it. Here is what Context Mode does for each, what we measured, and what we have not measured yet.

## The short version

- Fewer tokens is not a smaller bill. On Claude Opus 5.5 a cache write costs 40 times a cache read, so a change that breaks the cache can cost more than it saves.
- Context Mode trims long tool descriptions but keeps every tool, and the trim gives the same bytes on every turn. In one recording, a warm turn cost 42.5% less.
- The biggest leak in our own data was the cached prefix changing when it did not need to. The biggest cause was ours, and the fix is on by default.
- On Claude Haiku 5.5, a request above 100,000 prompt tokens pays 5 times for every token. Context Mode now books that tier drop as its own saving.
- Codex now gets the same tool trim, and its cost is measured per request.

## The problem: the bill does not follow the token count

A coding agent sends the whole conversation on every call. Prompt caching makes most of that cheap. On Claude Opus 5.5 the list price per million tokens is:

| Token type | Price per million tokens |
| --- | --- |
| Fresh input | $4 |
| 1-hour cache write | $8 |
| Cache read | $0.20 |

A read costs a twentieth of a fresh token. A write costs forty times a read. That ratio decides whether a change saves money: if it makes the provider write the prefix again, it has to remove a lot of read tokens to pay that write back. In our August 2026 benchmark grid on Claude Haiku 4.5 (write $2.00, read $0.10, so 20:1), cache writes were 45 to 94% of each cell's cost.

## How we learned it: tool search

On 2026-10-01 we tested tool search: send none of the 60 MCP tool definitions up front, and let the model search for the ones it needs. Prompt tokens per session fell 39%, from 176,196 to 107,495. Every task still passed, 12 of 12 in both arms.

Cost rose 149.6% over 12 paired runs on Opus 5.5.

**Tool search off** *$0.1043* **Tool search on, 39% fewer tokens** *$0.2107*

Mean cost per run, 12 paired runs on Opus 5.5, 60 MCP tools. Cost rose 149.6% (95% interval +104.1 to +195.1%).

The 60 definitions had been cheap cache reads. Taking them out caused new cache writes and uncached input. We did not ship it. Everything below cuts tokens without breaking the cache.

## Fix 1: trim tool schemas, keep every tool

Tool schemas sit at the head of every request. In one recording, Claude Code sent a 29-tool manifest of 103,253 bytes.

Context Mode does not remove tools. Each tool of 1,500 bytes or more keeps its name, its top-level parameters and its required list. Its description is cut at a sentence end at or before 420 characters, with a note that tells the model to call `context-mode-tool-schema` for the rest. The slim form depends only on the tool list, so it is the same bytes on every turn.

We measured it with a recording proxy in front of a real `claude -p` on a Haiku model, 7 turns:

|  | Full manifest | Slim manifest |
| --- | --- | --- |
| Manifest size | 103,253 B | 27,436 B |
| Prompt tokens, warm turn | 40,682 | 20,748 |
| Cost per warm turn | $0.004596 | $0.002640 (-42.5%) |
| The turn the manifest changes |  | $0.040917 |

Look at the last row. The turn where the manifest changes costs 8.9 times a warm turn, because the provider writes 20,143 tokens to cache again. That one turn takes about 20 warm turns to pay back. A trim that changed its output from turn to turn would never pay back. That is why the trim must give the same bytes every time.

The same study tried the other way first: removing 26 tools instead of trimming them made that turn 4.7 times more expensive.

### Now on Codex too

Codex 0.156.1 declares its tools in a different shape, and until now Context Mode did not trim them. It does now. The Codex trim cuts descriptions only:

- `exec` keeps its grammar, its rules and every nested tool signature, and goes from 14,650 to 10,273 bytes.
- `spawn_agent` keeps every parameter and goes from 2,726 to 1,131 bytes.
- Strict tools and Codex's built-in tools pass through unchanged.

On one captured gpt-6-astra request, with two live requests, input fell from 15,347 to 14,083 tokens: 1,264 fewer per request. In a live check, real `codex exec` turns that each had to call a tool passed 5 of 5 on gpt-6-astra and 3 of 3 on gpt-5.5.

Keep the size in proportion. After the first request, those 1,264 tokens are cache reads, so the dollar effect per turn is small. Context Mode also adds its own tokens to a Codex request: on 2026-09-25 we measured 1,799 more input tokens on one forwarded request, with memory and skills off. That was a different capture, so do not subtract one number from the other.

## Fix 2: hold the prefix still

In our own data, this mattered most.

In the 7 days to 2026-10-01, one heavy session spent $657.22 on its main agent. Of that, $331.30 (50.4%) was cold cache writes: requests that wrote their prompt to cache again. Almost every request used the 1-hour cache, yet 54 of the 65 cold writes came less than an hour after the previous request. The cache did not expire. Something changed the prefix.

So we found the cause of every cold write that came before the cache could expire, across three tenants, from 2026-09-22 to 2026-10-02: 198 cold writes. These are the causes we can name, plus the unknown ones:

| Cause | Cold writes |
| --- | --- |
| Context Mode moved the point where it cuts old messages | 90 |
| The client changed model or cache TTL (for example `/model` ) | 18 |
| Unknown, all before the fix | 44 |

The largest cause was ours. When a long thread nears the context window, Context Mode cuts old messages to keep it under the limit, and that cut point kept moving. The fix holds the cut in place and moves it only in planned steps:

- Before the fix: 89 such breaks over 9,119 main requests.
- After the fix: 1 break over 209 main requests, and that one was a planned step. The sample is small.

The fix has been on by default since 2026-10-03. The rule we took from it: every change Context Mode makes to a request is a pure function of its input, made once, and the same bytes after.

## Fix 3: count the price tier (Claude Haiku 5.5)

Claude Haiku 5.5 prices a whole request at one of two tiers. At or below 100,000 prompt tokens: $0.10 input, $0.50 output, $0.01 cache read, $0.20 1-hour cache write. Above that, every rate is 5 times higher. The prompt length counts all input tokens, cache reads and writes included.

So a request of 100,001 tokens pays 5 times for every token, the first 100,000 too. If Context Saving shortens a request from above the line to below it, the lower tier is a saving on top of the tokens removed.

What shipped on 2026-10-09: the ledger prices each call at its own tier and books the tier drop as its own row, "Lower price tier". In a unit test, a request folded from 150,003 to 60,003 tokens is priced at the low tier, and the tier drop alone is 4 times the bill. That is a test case, not a live measurement.

What did not ship: Context Mode does not aim at the line. It does not fold more because a request is near 100,000 tokens. We have no Haiku 5.5 runs through it yet.

## Fix 4: count Codex cost per request

You cannot cut what you do not count. Codex cost is now measured per request:

- **Reasoning tokens.** OpenAI bills reasoning inside output tokens. The usage row now keeps reasoning tokens apart. Cost does not change: reasoning is priced once, inside output.
- **ChatGPT login.** A ChatGPT plan does not pay per token. Context Mode prices those tokens at the API list price, and the Savings page labels them as equivalent API cost.
- **One dated price table**, read from OpenAI's pricing page on 2026-10-09. A model with no listed price (gpt-5.5, gpt-5.3) is marked unknown and left out of totals, not priced as a similar model.
- **A lower bound for the saving.** Removed pieces that can be counted exactly are priced. Changes to the request head are not, because OpenAI does not publish how it renders tools and instructions.

## What is not measured yet

- A paired cost A/B on Codex. It has not run.
- Haiku 5.5 through Context Mode. The tier saving comes from unit tests only.
- The Codex schema trim beyond one captured request and 8 live turns.
- The Claude Code schema trim A/B is a 7-turn recording on a Haiku model.
- The cause of 44 cold writes from before the prefix fix.

## How to try it

Context Mode sits between your coding agent and the model. It works with Claude Code and Codex today; OpenCode and Pi are coming soon. Your own Anthropic or OpenAI account pays the model.

- Set it up: [Quick start](https://context-mode.com/docs/quick-start).
- See how the trims and the ledger work: [How saving works](https://context-mode.com/docs/how-saving-works) and [Prompt caching](https://context-mode.com/docs/prompt-caching).
- The feature behind these savings: [Context Saving](https://context-mode.com/context-saving).
- What each plan includes: [Plans and limits](https://context-mode.com/docs/plans-and-limits) or [Pricing](https://context-mode.com/pricing).

[Start free](https://context-mode.com/docs/quick-start)[All posts](https://context-mode.com/blog)
