Context Saving · Thinking in Code for Claude Code and Codex

The script reads the data. The model reads what it prints.

With Thinking in Code, Claude Code or Codex writes a short script, the gateway runs it, and only what the script prints reaches the model. In one case on this page, the script read 7,011,585 bytes and the model read 407.

64%
lower cost on data analysis, Claude Code, Haiku 4.5
10%
lower cost on data analysis, Claude Code, Opus 5.5
Equal
correct answers in both arms, on both models

Paired runs, Claude Code 2.1.283, 2026-09-29. Codex: not measured yet. Every benchmark →

Key facts

What it is
A Context Mode Gateway feature: the agent writes a short script, the gateway runs it, and only what it prints reaches the model.
Who it is for
Agents that read large files, logs, API responses or many documents.
Clients
Claude Code and Codex. Public data runs in our sandbox; local files run on your machine.
Measured result
Long-context data analysis with Claude Code, 2026-09-29: 64% lower cost on Claude Haiku 4.5 and 10% lower on Claude Opus 5.5, with equal correct answers in both arms. Benchmarks.
Limits
A model that already filters its own data saves less, as Opus 5.5 shows. Codex is not measured yet.

Public data runs in our sandbox. Local scripts run on your machine.

run-codecontext-mode-run-code runs JavaScript in a Cloudflare Worker isolate, run by the gateway, so your agent's permission rules and hooks never see it. Egress is allow-listed, and every host it calls is checked against your Cage rules when Cage is on.
run-localcontext-mode-run-local runs through your agent's own shell, so it can read local files and localhost. The script runs on your machine. Only what it prints goes through the gateway to the model, and the archive keeps a copy.
print budgetLocal output over 16 KB keeps its first 6 KB and last 2 KB, a line that counts the rest, and a pointer to the full text.

The script reads 7,011,585 bytes. The model reads 407.

The data-analysis benchmark asks how many versions the react package lists, its dist-tags, and the dates of its first and latest stable release. The registry document is 7,011,585 bytes and lists 2,959 versions.

fetch toolGoing direct, the client's fetch tool hands the first 100,000 characters to a helper model, about 1.4% of the document.
shell scriptThe agent downloads and parses in its shell. Each command and its output stay in the conversation for every later turn.
run-codeThe script reads the whole document in the sandbox and returns 407 bytes: the count, the tags and two releases with their dates. The script is on the docs page.

Cloudflare and Anthropic described the pattern, and we run it for Claude Code and Codex.

Cloudflare's Code Mode and Anthropic's post on code execution with MCP both have the model write code against its tools, run it in a sandbox and read back only the result. Both are for teams building their own agent.

Context Mode gives the agents your team already uses the same tools. The gateway checks each network call against Cage and archives long output.

The full output is archived

The full output is archived before it is shortened. The agent reads any part back with context-mode-search and the pointer's id. Secrets and e-mail addresses are masked in the archive.


Measured

Against Claude Code going direct.

Questions about large npm registry documents, answered with a script, in paired sessions with Claude Code. Cost is list price applied to the token counts Anthropic returned for each call, including the extra model calls the gateway makes when it runs the code. All Context Saving mechanisms were on; this task is where Thinking in Code does most of the work.

Why the models differ: in our earlier runs of this task (2026-09-28), Opus 5.5 going direct often wrote its own download and script, so less raw data reached its prompt. Haiku 4.5 read more of the documents. Thinking in Code pays most where a model would otherwise read the data.

Paired runs, Claude Code 2.1.283, 2026-09-29, claude.ai login, cost at list price. Codex: not measured yet. How it works, how the model decides and why it is safe → · Every benchmark, with its method →


FAQ

Is this the same as code mode?

It is the same idea: the model writes code that reads the data, and only the result enters the context. Context Mode runs it for Claude Code and Codex through the gateway.

Where does the script run?

Public data runs in our sandbox. Local files run on your machine, and only what the script prints goes through the gateway.

Does it save the same on every model?

No. On data analysis, Haiku 4.5 cost 64% less and Opus 5.5 cost 10% less.

Keep raw data out of the prompt.

npx @context-mode/cli points Claude Code and Codex at the gateway.