Why Claude Code uses so many tokens and how to cut them
A coding agent pays for every byte a tool prints, again on each later turn, and for its tool schemas on every request. Context Mode Gateway folds tool output the first time it passes, keeps it the same on later turns so the cache keeps reading it, and archives the full output.
Key facts
- What it is
- The part of Context Mode Gateway that decides what each request carries: it folds tool output, runs data work as code, trims tool schemas and places cache marks.
- Who it is for
- Developers whose agents run tests, builds and installs, read long logs and large files, or start many subagents.
- Clients
- Claude Code and Codex. Tool-schema trimming runs for Claude Code only so far.
- Measured result
- On the 17 results we publish, with Claude Code on Claude Opus 5.5 and Claude Haiku 4.5, run on 2026-09-29, cost was 10 to 64% lower, with equal task success in both arms. Evaluation.
- Limits
- A short question with few tool calls and little output leaves little to save. Codex cost is not measured yet.
Summary
Context Saving is the part of the Context Mode gateway that decides what each request to the model carries. It shortens tool output the first time the output passes, runs data work as code so only the answer enters the prompt, and places cache marks so repeated text is read from the provider's cache. On the 17 results we publish, on Claude Opus 5.5 and Claude Haiku 4.5 with Claude Code, cost was 10 to 64% lower, and task success was equal in both arms on every one.
Paired A/B runs, Claude Code 2.1.283, 2026-09-29. Cost is list price applied to the token counts Anthropic returned for each call, hidden calls included; the runs used a claude.ai login, which has no per-token invoice (see Cost basis). "Likely" is the 95% interval. Codex cost is not measured yet. Full tables: Evaluation and Benchmarks.
Where an agent's tokens go
A coding agent sends the whole conversation to the model on every call. Four things decide the cost.
- Tool output stays. A test run, a log or a JSON file the agent reads enters the conversation and is sent again on every later call of the session. Claude Code saves Bash output over 30,000 characters to a file and shows a preview. Output under that line enters the prompt whole.
- Re-reads. The agent often reads the same thing twice. For WebFetch, Claude Code hands each fetched page to a small helper call with no cache mark, so each read of the page is charged as fresh input.
- Cache rewrites. The provider caches the start of the prompt. If a tool changes text inside that cached part, the next call writes it to cache again.
- Start cost. Tool schemas and the system prompt sit at the start of every request, and each new session writes part of that start to cache again.
Cache prices set the rules. Anthropic list prices per million tokens:
| Model | Input | Cache write, 1 hour | Cache read | Output |
|---|---|---|---|---|
| Claude Opus 5.5 | $4 | $8 | $0.20 | $20 |
| Claude Haiku 4.5 | $1 | $2 | $0.10 | $5 |
A cached token read back costs a twentieth of a fresh one on Opus 5.5 and a tenth on Haiku 4.5. A token written to cache costs 1.25 times a fresh one with a 5-minute mark ($1.25 on Haiku 4.5), and twice with a 1-hour mark. So a saving that rewrites cached text can cost more than it saves. Every mechanism below is built around that fact.
How it works
The gateway sits on the request path between your agent and the model API. It acts at three points. It does not change your files or your agent's settings. It does change what each request carries: it adds its own tools and a short context block, shortens tool schemas and tool output, and can answer a shell call itself. Each change is listed in Every change to a request.
Each mechanism below has its public name, what it does, what it never does, and the published result where its work shows. Every result ran with all mechanisms on, so no result isolates one mechanism.
Every change to a request
| Change | When |
|---|---|
| Adds its own tools: run-code, run-local, search, recall and the full-schema lookup, among others. | Every request. |
| Adds a short fixed block at the head of the first message: your stated preferences, your output style and your skill list. | Every request. Fixed for a conversation, so it is read from cache. |
| Adds memory from your other sessions. | A conversation's first request, a subagent's task, or the first request after the client replaces the conversation with a summary. |
| Cuts tool descriptions of 1,500 bytes or more at a sentence end at or before 420 characters. | Every request, Claude Code. |
| Adds one sentence to the shell tool's description, asking the model to report how many bytes a script read but did not print. | Every request. |
| Defers the search, file-pattern, web fetch, web search and notebook-read tools, or leaves them out. | Deferred for clients that load tools on demand; left out for clients that do not. |
| Moves text and cache marks in the system prompt and tool list, without changing a word. | Every request, Claude Code. |
| Shortens a tool result and archives the full text. | The first time the result passes. |
| Splits a WebFetch page from its question and gives the page a 5-minute cache mark. | WebFetch helper calls. |
| Runs helper calls on Claude Haiku 4.5. | Helper calls through a claude.ai login, as Claude Code does going direct. |
| Declines a shell call once and asks for run-local (steering). | A read-only command expected to print a lot. On by default. |
| Moves old history to the archive, with a pointer. | Very large threads, at fixed steps. |
The first-sight rule
The gateway changes a tool result only the first time the result passes. The change is a pure function of the result's text, so every later request carries the same bytes and the provider reads them from cache. It rewrites cached text only at fixed steps on very long threads: at each step, one request writes the history to cache again, and other aged changes ride that same request. A message that carries signed model thinking is never changed, because the provider rejects a turn whose signed thinking was edited.
Output folding
Shell output is mostly noise: passing tests, progress lines, repeats. The gateway runs 15 filters, in a fixed order, on the results of shell and search tools (Bash and Grep in Claude Code, exec and exec_command in Codex, and context-mode-run-local).
- Terminal cleanup. Progress redraws show only their last frame. Colour codes are removed.
- Folding. Passing tests fold into one counted line. Install, runner and browser-test progress fold the same way.
- Stack traces. A run of two or more library frames (
node_modules, the Python standard library and site packages, and similar) becomes one counted line. Your own frames and the error message stay. - Grouping. Repeated lines are counted once. Search matches, build errors and file listings are grouped by file. Long
git logoutput is grouped and counted. - JSON. A result that is one JSON document over 8,192 characters keeps every object key. Each array over 20 items keeps its first 3 items and its last 1, whatever they hold, and a line counts the rest.
- run-local print budget.
context-mode-run-localoutput over 16 KB keeps its first 6 KB and last 2 KB, whole lines only, and a line that counts the rest.
Test and build filters keep failing tests, error lines and summary lines as printed. A line no filter recognises stays as printed. The JSON filter and the run-local print budget cut by position, not by content, and stack traces lose library frames; the label counts what each filter removed, and the archive keeps all of it. Results under 1,024 characters pass through, and a fold is kept only if it saves at least 768 characters and 10% of the result. Read results and tool-server (MCP) results never go through folding. The shortened result ends with a label in this form (from the 600-test fixture below):
… 599 passing tests folded (<n> lines)
[context-mode compact — 599 passing tests folded — full output: context-mode-search {"docid":"context-mode-toolcap-<hash>-<length>"} (<length> B archived; every other line is as printed; do NOT re-run the command)]
The failing test, its assertion and the summary line stay as printed.
Measured on the test fixtures of the folding A/B, Claude Haiku 4.5, 2026-09-26.
Published results where its work shows: Test-Driven Debugging (Haiku 4.5 53%, Opus 5.5 22%), Large Tool-Output Handling (Opus 5.5 22% on a log, 21% on JSON) and Agentic Coding (Opus 5.5 29% on code search, 25% on a git log question). See Evaluation. Product page: Context Saving.
The current-turn cap
A single result over 24,576 characters, such as a very large file read or tool-server payload, is archived whole. The model gets its first 4,096 characters, its last 2,048 characters and a label that says the middle was cut and names the archive id. Error results are never capped. Claude Code already saves Bash output over 30,000 characters to a file, so in practice this cap applies to file reads and tool-server results.
Tool-schema slimming
Tool schemas sit at the start of every request. Each tool of 1,500 bytes or more keeps its name, its top-level parameters and its required list. Its description is cut at a sentence end at or before 420 characters, and a note tells the model how to get the rest. When the model needs the full definition, it calls context-mode-tool-schema, which answers from the schemas your agent sent on the same request. Tool order and cache marks are kept, and the slim form depends only on the tool list, so it is the same bytes on every turn.
Tools that return bulk (Grep, Glob, WebFetch, WebSearch, NotebookRead) are deferred for clients that load tools on demand, and left out for clients that do not. Read is never deferred or left out. This runs for Claude Code. It does not run for Codex yet.
Start-cost placement
Five transforms move text and cache marks so that a new session reads the start of its prompt from cache instead of writing it again. They never add, remove or reword text.
- System split. Claude Code joins the fixed and the per-session parts of its system prompt into one block. The gateway cuts it at the first per-session heading, so the fixed part gets its own cache mark.
- Lists move to their tools. The client's skill list and agent-type list move, byte for byte, into the descriptions of the tools they describe, which sit in the cached start.
- Lists last. The tools that carry a list move to the end of the tool array, so a changed list does not break the cache for the other tools.
- Trailing marks. A cache mark on a trailing system message moves to the message before it, so the conversation before it is read back.
- Head pin. The gateway's own block at the start of the prompt is fixed for a conversation and renewed only after 65 minutes idle, so a deploy of the gateway does not change it.
The provider accepts at most 4 cache marks (3 on some requests). The gateway never removes a mark your client set, and it adds its own only while a slot is free. If the provider still refuses the mark count, the request is retried once with no gateway marks. Transforms 1 to 4 are for Claude Code; OpenAI's cache has no marks.
Where its work shows: on the six short Opus 5.5 tasks (88 pairs), the input cost of a session's first request was 46% lower than going direct, likely 44 to 48%. No tool output exists yet on a first request, so this is the work of start-cost placement and schema slimming. The public evidence files hold per-pair totals; this figure comes from the per-request records of the same run.
The page cache
A WebFetch helper call sends the whole page with no cache mark, so a second question about the same page pays for the page again. The gateway splits that text into the page and the question, and marks the page for a 5-minute cache. Joined, the two parts are the client's text, byte for byte. The next read of that page within 5 minutes is a cache read. The first read writes the page to cache at the 5-minute price, 1.25 times a fresh read.
Memory and the fixed head block
A short fixed block starts every request: your stated preferences, your output style and your skill list. It is decided on a conversation's first request and kept byte for byte (see Head pin above), so later requests read it from cache. Other memory, such as facts and past sessions found by search, is added only on a conversation's first request, a subagent's task, or the first request after the client replaces the conversation with a summary. Later turns carry none of it, so they keep reading from cache. The model can still ask for memory with context-mode-recall or context-mode-search. See Memory.
Long sessions
- Long inputs. On any request, a long shell command or script (512 characters or more, in a tool call of 2,048 characters or more) becomes a pointer with its first and last 5 lines, the first time it passes. File edits and the file text they carry are never touched.
- Old results move to the archive. Past about 750,000 tokens, old tool results of 2,048 characters or more become a one-line pointer with a short summary. In the same pass, a repeated identical error becomes one line that names the first copy and the attempt count; the newest error of each tool stays whole. The boundary moves only at fixed steps, so the cache holds between steps.
- Whole old messages. Past about 800,000 tokens, or when a thread passes the model's own window, whole old messages, older user turns included, move to the archive with a pointer. Without this, the provider refuses the request. The first user turn always stays.
- Pre-heal. When the provider refuses a thread as too long, the gateway reads the limit off the refusal, cuts at a fixed boundary and keeps that plan for the thread. The next request is cut before it is sent, so it goes through on the first try and the kept part still reads from cache. On a stated 1M-token window, the plan is written from 60% of the window, before any refusal.
- Classifier fit. Claude Code's auto-mode safety check sends the whole transcript in one call and can pass the window. The gateway fits it by moving the oldest tool outputs to the archive first, then tool calls, then agent text. It never removes user text or the newest entries. A call that fits is sent unchanged.
Steering
Steering is on by default. Before a read-only shell command runs, the gateway looks at the output size that command family printed before. It steers only when the expected saving is more than twice the cost of one extra model call. It then declines the call once, with a short answer that starts context-mode: not run., and asks the agent to use context-mode-run-local and print only what it needs. The same command sent again goes through, so a model that needs the whole output gets it.
The archive and read-back
Before any output is shortened, the full text goes to your archive under an id. For tool output the id starts with context-mode-toolcap-. The short version names the id. The agent calls context-mode-search with that id, and an offset for later pages, and reads the text back. Secrets and e-mail addresses are masked before storage. See Search.
Parity with going direct
Two mechanisms do not save against going direct. They keep a request through Context Mode as cheap as the same request sent direct.
- Helper model. Claude Code makes small helper calls: WebFetch page reads, titles, summaries. Going direct, or with an Anthropic API key, it runs them on Claude Haiku 4.5. Through any custom base URL with a claude.ai login, Claude Code 2.1.283 runs them on the main model. The gateway recognises a helper by its exact shape and runs it on Haiku 4.5, as Claude Code does going direct. It never reroutes a call that carries a conversation.
- Side calls. A helper call with no tools gets no memory, skills or style block. A real conversation is never treated as a side call.
Where you see the saving
- Terminal. A banner line that starts with
context-mode ·gives the tokens the turn saved and the bytes it kept out. - Console. The Live page at console.context-mode.com shows running sessions and the saving per minute.
The console's ledger prices each request twice: at list price for the usage the provider returned, and for what Claude Code going direct would have sent, counted with the provider's own token counter. The saving is the difference between the two. A request gets a dollar figure only when its count is within 5% of the provider's prompt count.
Thinking in Code, in depth
Thinking in Code gives the model two tools. The model writes a short program, the program reads the data, and only what the program returns enters the conversation.
Bash in Claude Code, exec in Codex. The program runs on your machine, as a normal shell command. What it prints goes through the gateway to the model, and long output is archived. For local files and localhost.Before and after
The Data Analysis benchmark asks about the npm registry documents for react and express. Below is one question, measured by hand with the script under it on 2026-09-29 (UTC). It is one query, not a benchmark figure.
Claude Code's fetch tool hands the first 100,000 characters of a page to a helper model: about 1.4% of the 7,011,585-byte react document (2,959 versions). Sizes change as new versions ship.
The version count, the dist-tags, and the first and latest stable release with their dates, computed over the whole document.
// Node 18 or later, or context-mode-run-code. Prints the bytes in, the bytes out and the answer.
const r = await fetch("https://registry.npmjs.org/react");
const t = await r.text(), d = JSON.parse(t);
const stable = Object.keys(d.versions).filter((v) => !v.includes("-"))
.sort((a, b) => new Date(d.time[a]) - new Date(d.time[b]));
const latest = d["dist-tags"].latest;
const out = JSON.stringify({ versions: Object.keys(d.versions).length, dist_tags: d["dist-tags"],
first_stable: { version: stable[0], date: d.time[stable[0]] },
latest_stable: { version: latest, date: d.time[latest] } });
const bytes = (x) => new TextEncoder().encode(x).length;
console.log(bytes(t), "bytes in,", bytes(out), "bytes out");
console.log(out);
Without Thinking in Code, the agent has two choices. It can read the page through its fetch tool, which sees only the first part of the document. Or it can write its own download and script in the shell, where every command and its output stay in the conversation. With Thinking in Code, the program reads the whole document outside the conversation and returns only the answer.
How the model decides
The gateway adds the two tools to each request. Each tool description says when to use it and when not: run-code for a public page or API that has to be fetched and parsed, or for work that can be computed instead of read; run-local for files on your machine and localhost. Nothing forces the tool, with one exception: when steering expects a large output, it declines the shell call once and asks for run-local (see Steering). Otherwise the model picks the tool when the task fits, the same way it picks any other tool.
Why it is safe
- Sandboxed. run-code runs in an isolate with no access to your machine. Egress is allow-listed, and each network call is checked against Cage before it leaves when Cage is on.
- Outside your agent's rules. The gateway runs run-code itself, so Claude Code and Codex permission rules and hooks never see that call. Cage is the only rule set that checks it.
- Your rules apply to run-local. run-local becomes your agent's own shell call, so your agent's permission rules and hooks apply to it as to any other shell command.
- Only the return value enters. A return value over 60,000 characters is archived in full; the model gets its first 30,000 and last 20,000 characters and the archive id. Printed run-local output over 16 KB keeps its first 6 KB and last 2 KB, whole lines only, and a line that counts the rest.
- Nothing is dropped. The full output is archived before the view is shortened, and
context-mode-searchreads it back by id.
Measured value
All Context Saving mechanisms were on in these runs. This is the task where Thinking in Code does most of the work.
| Long-Context Data Analysis | Saved | Direct | Context Mode | Per 1,000 tasks | Head-to-head | Correct answers |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | 64%Haiku 4.5, likely 52 to 77% | $0.175 | $0.062 | $113 | 12 of 12 | 11 of 12 in each arm |
| Claude Opus 5.5 | 10%Opus 5.5, likely 4 to 15% | $0.223 | $0.201 | $22 | 15 of 20 | 20 of 20 in each arm |
Business value. On Haiku 4.5 a data-analysis task cost $0.062 through Context Mode against $0.175 direct: $113 less per 1,000 tasks. On Opus 5.5 it was $22 less per 1,000 tasks. The Context Mode figures include the extra model calls the gateway makes when it answers run-code on its side.
Why the two models differ. In our earlier runs of this task (2026-09-28), Opus 5.5 going direct often wrote its own download and script, so less raw data reached its prompt and the gap is smaller. Haiku 4.5 going direct read more of the documents into its prompt. Thinking in Code pays most where a model would otherwise read the data.
The pattern is public. Anthropic's engineering post on code execution with MCP and Cloudflare's Code Mode post both describe a model writing code against its tools and reading back only the result. Context Mode runs it on the request path of the agents your team already uses, with Cage checks and a full archive. Product page: Thinking in Code.
Try it in 5 minutes
You run a big test suite and ask one question about it. The agent gets the answer. The full test log stays out of the conversation.
Setup
npx @context-mode/cli # sign in and connect Claude Code, then restart Claude Code
git clone https://github.com/expressjs/express.git && cd express && npm install
claude
Prompt
Run the full test suite and tell me if anything fails.
What you will see
All tests pass — 1261 tests total.
⏺ context-mode · context 67 KB lighter than your raw transcript · kept 71 KB out of what the model read
- The answer.
npm testprints one line per test, about 70 KB. The count matches the last line mocha prints. - kept 71 KB out. The gateway kept the full log out of what the model read and gave it the part it needed. If the agent asks for an exact line later, the full log is still there to look up.
- Your numbers will differ. In a second run the line said
kept 65 KB out. When the agent filters the log itself first (for example withgrep), there is less to keep out: in a run with two failing tests the line saidkept 5 KB out, and the agent still named both tests and the cause. - In the console. The Value page shows what was kept out, turn by turn.
We ran this on 2026-10-04 in real Claude Code sessions (claude -p 2.1.283, Claude Haiku 4.5) through the gateway, each on a new test account. Three runs. All three gave the right test count, or the right failing tests.
Why it matters. Test runs, builds and logs are the biggest things an agent reads in a day. The model gets the answer, and the context stays small for the rest of the session.
Compared with other approaches
Each approach below solves part of the problem. The table says what each one does and does not do, as we understand it.
| Approach | What it does | What it does not do |
|---|---|---|
| Truncation and preview in the client | Cuts a large output or saves it to a file and shows a preview. | Output under the cut line enters whole and stays for the session. What was cut is not in the conversation, and the preview does not say which lines matter. |
| Summarising old turns | Replaces a long conversation with a written summary to free room. | The summary is lossy: exact output, ids and numbers can go. The next request writes the new start to cache again. |
| RAG and indexing | Stores documents and returns the passages that match a query. | It does not shrink tool output that is already in the conversation. The model still needs the right query, and a passage is not a computed answer. |
| Output filters on your machine | Shell wrappers or hooks shorten command output before the agent sees it. | They run per machine and per client. Unless they keep a copy, what they cut is gone. They cannot move cache marks or fit requests to the window. |
| Provider prompt caching alone | Reads a repeated prompt start at a lower price. | It does not make the prompt smaller. It only pays when the cached text stays the same, so any edit of earlier text is a new cache write. |
| Context Mode | Acts on the request itself: first-sight folding with a full archive, code runs that return only the answer, cache-mark placement, and window fitting. | It does not edit your files or your agent's settings. It rewrites cached text only at fixed steps on very long threads, one request per step. |
Evaluation
Method
- Paired A/B. Each pair runs the same scripted task twice, back to back: once with Claude Code direct to Anthropic, once through Context Mode. The order alternates every pair. Pairing cancels drift during the run, such as provider load and cache state.
- One task per row. Each row is one scripted task, repeated in every pair. The task is named in the row.
- The real client. Claude Code 2.1.283 (
claude -p) in both arms, fresh config, no plugin, for every run. - The login. All runs used a claude.ai login. With an Anthropic API key, the helper model and cache-write prices differ; we have not measured that setup.
- What the Context Mode arm carried. Memory, skills and an output-style block were on in every Context Mode run. The style block asks for short answers and follows the login we used. Without it, the saving can differ.
- Written down first. The tasks, the number of pairs and the analysis for each run were committed before its first pair. One warm-up pair per task ran first and does not count.
- Gateway builds. The published rows ran on three gateway builds, all on 2026-09-29; each evidence file names its build. When we fixed a defect, we ran the row again on the new build, pre-registered the same way, and the new run replaced the old one.
- Deploy guard. The runner reads the deployed gateway version before and after each pair. If it changed, the pair runs again.
- Cost. List price applied to the token counts Anthropic returned for each call. See Cost basis.
- No pair dropped. Every valid pair counts. The saving is the mean difference over the mean direct cost, with a 95% t interval.
- What counts. We publish a result only when its whole 95% interval is below zero. Each task has a success check (tests pass, the right answer), and success was equal in both arms on every result below.
Results by category
| Category and task | Saved | Direct | Context Mode | Per 1,000 tasks | Head-to-head | Evidence |
|---|---|---|---|---|---|---|
| Long-Context Data Analysis | ||||||
Questions about the react and express npm registry documents, answered with a script | 64%Haiku 4.5, likely 52 to 77% | $0.175 | $0.062 | $113 | 12 of 12 | JSON · Pre-registration |
Questions about the react and express npm registry documents, answered with a script | 10%Opus 5.5, likely 4 to 15%, later Gateway commit | $0.223 | $0.201 | $22 | 15 of 20 | JSON · Pre-registration |
| Test-Driven Debugging | ||||||
npm test fix loop: run the tests, fix the code, run them again | 53%Haiku 4.5, likely 49 to 56% | $0.069 | $0.032 | $36 | 12 of 12 | JSON · Pre-registration |
npm test fix loop, short | 22%Opus 5.5, likely 16 to 28% | $0.087 | $0.068 | $19 | 15 of 16 | JSON · Pre-registration |
npm test fix loop, long | 21%Opus 5.5, likely 14 to 28% | $0.085 | $0.067 | $18 | 11 of 12 | JSON · Pre-registration |
| Agentic Coding (short, one-prompt coding tasks) | ||||||
Code search (grep over 150 files) | 29%Opus 5.5, likely 23 to 34% | $0.093 | $0.066 | $27 | 16 of 16 | JSON · Pre-registration |
| One-line edit | 28%Opus 5.5, likely 14 to 42% | $0.075 | $0.054 | $21 | 11 of 12 | JSON · Pre-registration |
Repo history question (git log, 320 commits) | 25%Opus 5.5, likely 16 to 34% | $0.071 | $0.053 | $18 | 15 of 16 | JSON · Pre-registration |
Code search (grep) | 29%Haiku 4.5, likely 4 to 54% | $0.031 | $0.022 | $9 | 18 of 20 | JSON · Pre-registration |
| One-line edit | 20%Haiku 4.5, likely 15 to 25% | $0.025 | $0.020 | $5 | 8 of 8 | JSON · Pre-registration |
Repo history question (git log) | 19%Haiku 4.5, likely 16 to 23% | $0.025 | $0.020 | $5 | 12 of 12 | JSON · Pre-registration |
| Large Tool-Output Handling | ||||||
| A question about one large log | 22%Opus 5.5, likely 15 to 30% | $0.091 | $0.071 | $20 | 15 of 16 | JSON · Pre-registration |
| A question about one large JSON file | 21%Opus 5.5, likely 13 to 30% | $0.086 | $0.068 | $18 | 12 of 12 | JSON · Pre-registration |
| A question about one large JSON file | 16%Haiku 4.5, likely 9 to 23% | $0.032 | $0.027 | $5 | 18 of 20 | JSON · Pre-registration |
| Multi-Agent Orchestration | ||||||
| Three parallel subagents, then one more | 15%Opus 5.5, likely 12 to 19% | $0.481 | $0.408 | $73 | 20 of 20 | JSON · Pre-registration |
Cost at list price, paired runs, 2026-09-29, Claude Code 2.1.283. Cost per task is the mean cost. Per 1,000 tasks is the mean saving per task, times 1,000. Head-to-head counts the pairs where Context Mode cost less. We also ran the test-fix loop in Ruby to check the result is not tied to one toolchain: Opus 5.5 61% lower (likely 52 to 70%), Haiku 4.5 46% lower (likely 34 to 57%), 8 of 8 pairs won on each model.
Cost basis: provider usage at list price, not the client's estimate
Cost is list price applied to the token counts Anthropic returned for each call. The runs used a claude.ai login, which has no per-token invoice.
- Direct arm. Claude Code's own total. In 52 of 52 runs we checked, it equals its own per-model token counts at each model's list price.
- Context Mode arm. The gateway's per-call record of the same usage fields Anthropic returned, hidden calls included, at list price.
Through a gateway, Claude Code's own figure is wrong in two ways, so we do not use it for the Context Mode arm. It prices the WebFetch helper at the main model's price, although the call ran on Haiku 4.5. And it does not see the calls the gateway makes when it answers run-code on its side.
The check that the gateway record misses nothing: over the 224 valid Context Mode runs on gateway build 3bedf1ba, the gateway record holds at least every token Claude Code saw, in each class. In the 172 runs with no gateway-side call, the two are equal. On Data Analysis (Haiku 4.5, and an earlier Opus 5.5 run), in 32 of 32 runs the difference is exactly the gateway-side calls at list price. Pair 1 of that earlier Opus 5.5 Data Analysis run, call by call:
| Context Mode arm, pair 1 | Calls | At list price |
|---|---|---|
| Calls Claude Code made, on Claude Opus 5.5 (Claude Code's own figure) | 3 | $0.156003 |
| Calls the gateway made on its side, after it answered run-code | 2 | $0.050448 |
| Context Mode arm, all calls | 5 | $0.206451 |
| Direct arm, pair 1, Claude Code's own total | $0.235333 |
Rows are rounded to the sixth decimal. The figures for this pair are in the published evidence file: "Large documents processed with a script", pair 1, $0.2353328 direct and $0.206451 through Context Mode. Claude Code's own figure is lower than the gateway record because it misses the gateway-side calls; we publish the gateway record.
Safety and limits
What it keeps
- A full archive. Text is archived before it is shortened, and read back by id. Secrets and e-mail addresses are masked before storage.
- User text. Compaction, the caps and the first level of long-session archiving never shorten or rewrite what you typed. When a thread passes the model's window, or about 800,000 tokens, whole old messages, older user turns included, move to the archive with a pointer, and your first turn stays in place.
- Signed thinking. A message that carries signed model thinking is never changed.
- The cache. Every change is made at first sight or at a fixed step, and it produces the same bytes on every later request.
- Failures. Test and build filters keep failing tests, error lines and summary lines as printed. The JSON filter and the run-local print budget cut by position, whatever the lines hold, and stack traces lose library frames; the archive keeps all of it. Error results are never capped.
When the saving is small
- A short question with few tool calls and little output leaves little to save.
- A model that already filters its own data saves less from Thinking in Code, as Opus 5.5 shows on data analysis.
What it does not do
- It does not restore images. On threads past about 750,000 tokens, an old image is replaced by a note that it was removed.
- It does not edit your files or your agent's settings, and it does not resell tokens. Your provider bills you directly.
- We have not measured cost on Codex, with an Anthropic API key, or on models other than Claude Opus 5.5 and Claude Haiku 4.5. Tool-schema slimming does not run for Codex yet.
- Memory, skills and an output-style block were on in every Context Mode run, so the published results include them. The style block asks for short answers; without it, the saving can differ.
- Some results depend on how Claude Code 2.1.283 builds its requests. A new client release can change them, so each result names its client version.
- The console's ledger matched the paired runs on 2026-09-26 (Claude Haiku 4.5). We have not repeated that check on the 2026-09-29 benchmark.
Reproduce the numbers
Every result on this page comes from an evidence file on this site. Each file lists every pair with the cost per arm and the success check. The evidence index gives the method, the client version, the gateway version and what we removed (prompt text, session and tenant ids). This script recomputes a row from a file:
import json, statistics as st, sys
T = {7: 2.365, 11: 2.201, 15: 2.131, 19: 2.093} # t at 97.5%, by degrees of freedom
path, task = sys.argv[1], sys.argv[2] # e.g. final-haiku45.json W3
w = json.load(open(path))["workloads"]
w = w[task] if isinstance(w, dict) else next(x for x in w if x["id"] == task)
P = [p["plain"]["usd"] for p in w["pairs"]]
D = [p["gateway"]["usd"] - p["plain"]["usd"] for p in w["pairs"]]
n, mp, md = len(D), st.mean(P), st.mean(D)
h = T[n - 1] * st.stdev(D) / n ** 0.5
print(task, n, "pairs:", round(100 * md / mp, 1), "% (likely",
round(100 * (md + h) / mp, 1), "to", round(100 * (md - h) / mp, 1), "%)",
"won", sum(d < 0 for d in D))
Run on final-haiku45.json W3, it prints 12 pairs, −64.4%, likely −52.3 to −76.5%, won 12. On v4-opus55-W1-W2-W3-W4.json W2 it prints 20 pairs, −15.2%, likely −11.7 to −18.7%, won 20. Your own saving depends on your mix of work; the console measures it from your traffic.
FAQ for CTOs
Does a proxy break my prompt cache?
It is built not to. A result is changed only the first time it passes, and the same bytes go out on every later request. The gateway never removes your client's cache marks and never sends more than the provider allows. It rewrites cached text only at fixed steps on very long threads. The results are measured cost, so every cache write is in them.
Does quality drop?
Task success was equal in both arms on every published result, with a check per task (tests pass, the right answer). Test and build filters keep failures, errors and summaries as printed.
Is this truncation that loses data?
No. The full text is archived before the short view is made, and the short view names its archive id. The agent reads the text back with context-mode-search. Secrets and e-mail addresses are masked in the archive.
Did you pick the best benchmarks?
We publish tasks where the whole 95% interval is below zero. Each run's tasks, pair count and analysis were committed before its first pair, and no pair was dropped. Each result row links its pre-registration, with the commit and its date. The published rows ran on three gateway builds on 2026-09-29. When we fixed a defect, we ran the row again on the new build, pre-registered the same way, and the new run replaced the old one. Benchmarks lists the sources.
Aren't 8 to 20 pairs too few?
The design is paired, so each pair is its own control, and the interval is a t interval over the paired differences. Head-to-head counts are shown with each row: Multi-Agent won 20 of 20 pairs on Opus 5.5.
Is that your meter, not my bill?
The runs used a claude.ai login, so there is no per-token invoice to show. Both arms are list price applied to the token counts Anthropic returned. The direct arm is Claude Code's own total. The Context Mode arm is the gateway's per-call record, which holds at least every token Claude Code saw. See Cost basis.
Is this production data?
No. It is a controlled A/B with the real client (claude -p), not a replay. Your saving depends on your work; the console measures it from your own traffic.
Is a proxy in my path a risk?
Leaving takes one step: remove ANTHROPIC_BASE_URL from ~/.claude/settings.json; the CLI kept a backup. A change to the request shape must pass a real-client test through the gateway before it ships. See How do I leave?
What about privacy?
run-local runs on your machine; what it prints goes through the gateway to the model, and long output is archived. Each run-code network call is allow-listed and checked against Cage when Cage is on. Archives live on Cloudflare, one store per account. Prompts have no automatic expiry yet, and there is no one-click delete or export yet; ask us and we delete your data. See Data.
Won't the model provider ship this?
The code-execution pattern is public, and we credit it. The results on this page come from work on the request itself: Thinking in Code, first-sight folding with a full archive, cache-mark placement and page caching. A plugin inside the client cannot move cache marks or change the request the client sends. Some results depend on how Claude Code 2.1.283 builds its requests, so a client release can change them; each result names its client version.
Compared with other tools: see the landscape.