Prompt caching in Claude Code and Codex

The provider caches the start of the prompt, and on Claude Opus 5.5 a cached token read back costs a twentieth of a fresh one. If anything changes text inside that cached part, the next call writes it to cache again at a higher price, so Context Mode Gateway changes tool output only the first time it passes and sends the same bytes on every later turn.

Key facts

What it is
How the gateway keeps Claude Code and Codex requests cache-friendly: the first-sight rule, cache marks on the fixed part of the prompt, and a cache for fetched pages.
Who it is for
Developers and CTOs who see high cache-write costs in Claude Code, or want to know whether a proxy breaks the prompt cache.
Clients
Claude Code and Codex. The system split and the list moves act on Claude Code requests.
Measured result
The results are measured cost, so every cache write is in them: Claude Code on Claude Opus 5.5 cost 65.5% less on a long agentic coding session, which went from $7.20 to $2.49, 4 of 4 pairs cheaper, run on 2026-10-03. Over 168 shorter paired runs on 2026-09-29, 155 cost less and 168 of 168 tasks passed in both arms. Method and evidence.
Limits
On very long threads the gateway rewrites cached text at fixed steps, and one request writes the history to cache again at each step. The provider accepts at most 4 cache marks.

What cache reads and writes cost

A cached token read back costs a twentieth of a fresh one on Opus 5.5 and a tenth on Haiku 4.5. A token written to cache costs 1.25 times a fresh one with a 5-minute mark, and twice with a 1-hour mark (Anthropic prompt caching, read 2026-10-01). So a saving that rewrites cached text can cost more than it saves. OpenAI also prices repeated input far lower than new input (OpenAI prompt caching).

What makes a coding agent write the cache again

What the gateway does

First sightA tool result is changed only the first time it passes. The change depends only on the result's text, so every later request carries the same bytes and the provider reads them from cache.
System splitThe gateway cuts Claude Code's system prompt at the first per-session heading, so the fixed part gets its own cache mark.
Lists moveThe skill and agent-type lists move, unchanged, into the descriptions of the tools they describe, and those tools move to the end of the tool array. Nothing is added, removed or reworded.
Page cacheA WebFetch helper call is split into the page and the question, and the page is marked for a 5-minute cache. The next read of that page within 5 minutes is a cache read.
Mark budgetThe gateway never removes a mark your client set, and adds its own only while a slot is free. If the provider refuses the mark count, the request is retried once with no gateway marks.

Every change is listed, with its limits, in Context Saving.

FAQ

Does a proxy break my prompt cache?

Context Mode Gateway is built not to. A result is changed only the first time it passes, and the same bytes go out on every later request. It rewrites cached text only at fixed steps on very long threads.

What does a cache write cost?

1.25 times a fresh token with a 5-minute mark, and twice with a 1-hour mark, on Anthropic models.

Why does Claude Code pay full price for the same WebFetch page again?

Claude Code hands each fetched page to a helper call with no cache mark. The gateway marks the page for 5 minutes, so a second question about it within that time is a cache read.

Does the gateway change my cache settings?

No. It places cache marks on requests and never removes your client's marks. It does not change your provider account.