Context Saving · Thinking in Code for Claude Code and Codex
The script reads the data. The model reads what it prints.
With Thinking in Code, Claude Code or Codex writes a short script, the gateway runs it, and only what the script prints reaches the model. In one case on this page, the script read 7,011,585 bytes and the model read 407.
Paired runs, Claude Code 2.1.283, 2026-09-29. Codex: not measured yet. Every benchmark →
Key facts
- What it is
- A Context Mode Gateway feature: the agent writes a short script, the gateway runs it, and only what it prints reaches the model.
- Who it is for
- Agents that read large files, logs, API responses or many documents.
- Clients
- Claude Code and Codex. Public data runs in our sandbox; local files run on your machine.
- Measured result
- Long-context data analysis with Claude Code, 2026-09-29: 64% lower cost on Claude Haiku 4.5 and 10% lower on Claude Opus 5.5, with equal correct answers in both arms. Benchmarks.
- Limits
- A model that already filters its own data saves less, as Opus 5.5 shows. Codex is not measured yet.
The script reads 7,011,585 bytes. The model reads 407.
The data-analysis benchmark asks how many versions the react package lists, its dist-tags, and the dates of its first and latest stable release. The registry document is 7,011,585 bytes and lists 2,959 versions.
Cloudflare and Anthropic described the pattern, and we run it for Claude Code and Codex.
Cloudflare's Code Mode and Anthropic's post on code execution with MCP both have the model write code against its tools, run it in a sandbox and read back only the result. Both are for teams building their own agent.
Context Mode gives the agents your team already uses the same tools. The gateway checks each network call against Cage and archives long output.
The full output is archived before it is shortened. The agent reads any part back with context-mode-search and the pointer's id. Secrets and e-mail addresses are masked in the archive.
Measured
Against Claude Code going direct.
Questions about large npm registry documents, answered with a script, in paired sessions with Claude Code. Cost is list price applied to the token counts Anthropic returned for each call, including the extra model calls the gateway makes when it runs the code. All Context Saving mechanisms were on; this task is where Thinking in Code does most of the work.
Why the models differ: in our earlier runs of this task (2026-09-28), Opus 5.5 going direct often wrote its own download and script, so less raw data reached its prompt. Haiku 4.5 read more of the documents. Thinking in Code pays most where a model would otherwise read the data.
Paired runs, Claude Code 2.1.283, 2026-09-29, claude.ai login, cost at list price. Codex: not measured yet. How it works, how the model decides and why it is safe → · Every benchmark, with its method →
FAQ
Is this the same as code mode?
It is the same idea: the model writes code that reads the data, and only the result enters the context. Context Mode runs it for Claude Code and Codex through the gateway.
Where does the script run?
Public data runs in our sandbox. Local files run on your machine, and only what the script prints goes through the gateway.
Does it save the same on every model?
No. On data analysis, Haiku 4.5 cost 64% less and Opus 5.5 cost 10% less.
Keep raw data out of the prompt.
npx @context-mode/cli points Claude Code and Codex at the gateway.