Why Claude Code uses so many tokens and how to cut them

A coding agent pays for every byte a tool prints, again on each later turn, and for its tool schemas on every request. Context Mode Gateway folds tool output the first time it passes, keeps it the same on later turns so the cache keeps reading it, and archives the full output.

Key facts

What it is
The part of Context Mode Gateway that decides what each request carries: it folds tool output, runs data work as code, trims tool schemas and places cache marks.
Who it is for
Developers whose agents run tests, builds and installs, read long logs and large files, or start many subagents.
Clients
Claude Code and Codex. Tool-schema trimming runs for Claude Code only so far.
Measured result
On the 17 results we publish, with Claude Code on Claude Opus 5.5 and Claude Haiku 4.5, run on 2026-09-29, cost was 10 to 64% lower, with equal task success in both arms. Evaluation.
Limits
A short question with few tool calls and little output leaves little to save. Codex cost is not measured yet.

Summary

Context Saving is the part of the Context Mode gateway that decides what each request to the model carries. It shortens tool output the first time the output passes, runs data work as code so only the answer enters the prompt, and places cache marks so repeated text is read from the provider's cache. On the 17 results we publish, on Claude Opus 5.5 and Claude Haiku 4.5 with Claude Code, cost was 10 to 64% lower, and task success was equal in both arms on every one.

Paired A/B runs, Claude Code 2.1.283, 2026-09-29. Cost is list price applied to the token counts Anthropic returned for each call, hidden calls included; the runs used a claude.ai login, which has no per-token invoice (see Cost basis). "Likely" is the 95% interval. Codex cost is not measured yet. Full tables: Evaluation and Benchmarks.

Where an agent's tokens go

A coding agent sends the whole conversation to the model on every call. Four things decide the cost.

Cache prices set the rules. Anthropic list prices per million tokens:

ModelInputCache write, 1 hourCache readOutput
Claude Opus 5.5$4$8$0.20$20
Claude Haiku 4.5$1$2$0.10$5

A cached token read back costs a twentieth of a fresh one on Opus 5.5 and a tenth on Haiku 4.5. A token written to cache costs 1.25 times a fresh one with a 5-minute mark ($1.25 on Haiku 4.5), and twice with a 1-hour mark. So a saving that rewrites cached text can cost more than it saves. Every mechanism below is built around that fact.

How it works

The gateway sits on the request path between your agent and the model API. It acts at three points. It does not change your files or your agent's settings. It does change what each request carries: it adds its own tools and a short context block, shortens tool schemas and tool output, and can answer a shell call itself. Each change is listed in Every change to a request.

Your agent sends each request to the Context Mode gateway, which forwards it to the model API. The gateway acts on every request, before a tool runs, and when a tool result arrives. Long output goes to your archive. Your agent Claude Code or Codex Context Mode gateway Model API Anthropic or OpenAI Your archive full text, read back by id 1 every request: schemas, cache marks 2 before a tool runs: Cage, code 3 result arrives: folding, caps
Point 1 runs on every request. Point 2 runs when the model asks for a tool call. Point 3 runs when a tool result comes back, before the model reads it.

Each mechanism below has its public name, what it does, what it never does, and the published result where its work shows. Every result ran with all mechanisms on, so no result isolates one mechanism.

Every change to a request

ChangeWhen
Adds its own tools: run-code, run-local, search, recall and the full-schema lookup, among others.Every request.
Adds a short fixed block at the head of the first message: your stated preferences, your output style and your skill list.Every request. Fixed for a conversation, so it is read from cache.
Adds memory from your other sessions.A conversation's first request, a subagent's task, or the first request after the client replaces the conversation with a summary.
Cuts tool descriptions of 1,500 bytes or more at a sentence end at or before 420 characters.Every request, Claude Code.
Adds one sentence to the shell tool's description, asking the model to report how many bytes a script read but did not print.Every request.
Defers the search, file-pattern, web fetch, web search and notebook-read tools, or leaves them out.Deferred for clients that load tools on demand; left out for clients that do not.
Moves text and cache marks in the system prompt and tool list, without changing a word.Every request, Claude Code.
Shortens a tool result and archives the full text.The first time the result passes.
Splits a WebFetch page from its question and gives the page a 5-minute cache mark.WebFetch helper calls.
Runs helper calls on Claude Haiku 4.5.Helper calls through a claude.ai login, as Claude Code does going direct.
Declines a shell call once and asks for run-local (steering).A read-only command expected to print a lot. On by default.
Moves old history to the archive, with a pointer.Very large threads, at fixed steps.

The first-sight rule

The gateway changes a tool result only the first time the result passes. The change is a pure function of the result's text, so every later request carries the same bytes and the provider reads them from cache. It rewrites cached text only at fixed steps on very long threads: at each step, one request writes the history to cache again, and other aged changes ride that same request. A message that carries signed model thinking is never changed, because the provider rejects a turn whose signed thinking was edited.

Output folding

Shell output is mostly noise: passing tests, progress lines, repeats. The gateway runs 15 filters, in a fixed order, on the results of shell and search tools (Bash and Grep in Claude Code, exec and exec_command in Codex, and context-mode-run-local).

Test and build filters keep failing tests, error lines and summary lines as printed. A line no filter recognises stays as printed. The JSON filter and the run-local print budget cut by position, not by content, and stack traces lose library frames; the label counts what each filter removed, and the archive keeps all of it. Results under 1,024 characters pass through, and a fold is kept only if it saves at least 768 characters and 10% of the result. Read results and tool-server (MCP) results never go through folding. The shortened result ends with a label in this form (from the 600-test fixture below):

… 599 passing tests folded (<n> lines)
[context-mode compact — 599 passing tests folded — full output: context-mode-search {"docid":"context-mode-toolcap-<hash>-<length>"} (<length> B archived; every other line is as printed; do NOT re-run the command)]
npm test output, 600 tests, 1 failing
17.0 KB → 1.2 KB

The failing test, its assertion and the summary line stay as printed.

Ruby minitest -v, 401 tests, 1 failing
16.2 KB → 0.7 KB

Measured on the test fixtures of the folding A/B, Claude Haiku 4.5, 2026-09-26.

Published results where its work shows: Test-Driven Debugging (Haiku 4.5 53%, Opus 5.5 22%), Large Tool-Output Handling (Opus 5.5 22% on a log, 21% on JSON) and Agentic Coding (Opus 5.5 29% on code search, 25% on a git log question). See Evaluation. Product page: Context Saving.

The current-turn cap

A single result over 24,576 characters, such as a very large file read or tool-server payload, is archived whole. The model gets its first 4,096 characters, its last 2,048 characters and a label that says the middle was cut and names the archive id. Error results are never capped. Claude Code already saves Bash output over 30,000 characters to a file, so in practice this cap applies to file reads and tool-server results.

Tool-schema slimming

Tool schemas sit at the start of every request. Each tool of 1,500 bytes or more keeps its name, its top-level parameters and its required list. Its description is cut at a sentence end at or before 420 characters, and a note tells the model how to get the rest. When the model needs the full definition, it calls context-mode-tool-schema, which answers from the schemas your agent sent on the same request. Tool order and cache marks are kept, and the slim form depends only on the tool list, so it is the same bytes on every turn.

Tools that return bulk (Grep, Glob, WebFetch, WebSearch, NotebookRead) are deferred for clients that load tools on demand, and left out for clients that do not. Read is never deferred or left out. This runs for Claude Code. It does not run for Codex yet.

Start-cost placement

Five transforms move text and cache marks so that a new session reads the start of its prompt from cache instead of writing it again. They never add, remove or reword text.

The provider accepts at most 4 cache marks (3 on some requests). The gateway never removes a mark your client set, and it adds its own only while a slot is free. If the provider still refuses the mark count, the request is retried once with no gateway marks. Transforms 1 to 4 are for Claude Code; OpenAI's cache has no marks.

Where its work shows: on the six short Opus 5.5 tasks (88 pairs), the input cost of a session's first request was 46% lower than going direct, likely 44 to 48%. No tool output exists yet on a first request, so this is the work of start-cost placement and schema slimming. The public evidence files hold per-pair totals; this figure comes from the per-request records of the same run.

The page cache

A WebFetch helper call sends the whole page with no cache mark, so a second question about the same page pays for the page again. The gateway splits that text into the page and the question, and marks the page for a 5-minute cache. Joined, the two parts are the client's text, byte for byte. The next read of that page within 5 minutes is a cache read. The first read writes the page to cache at the 5-minute price, 1.25 times a fresh read.

Memory and the fixed head block

A short fixed block starts every request: your stated preferences, your output style and your skill list. It is decided on a conversation's first request and kept byte for byte (see Head pin above), so later requests read it from cache. Other memory, such as facts and past sessions found by search, is added only on a conversation's first request, a subagent's task, or the first request after the client replaces the conversation with a summary. Later turns carry none of it, so they keep reading from cache. The model can still ask for memory with context-mode-recall or context-mode-search. See Memory.

Long sessions

Steering

Steering is on by default. Before a read-only shell command runs, the gateway looks at the output size that command family printed before. It steers only when the expected saving is more than twice the cost of one extra model call. It then declines the call once, with a short answer that starts context-mode: not run., and asks the agent to use context-mode-run-local and print only what it needs. The same command sent again goes through, so a model that needs the whole output gets it.

The archive and read-back

Before any output is shortened, the full text goes to your archive under an id. For tool output the id starts with context-mode-toolcap-. The short version names the id. The agent calls context-mode-search with that id, and an offset for later pages, and reads the text back. Secrets and e-mail addresses are masked before storage. See Search.

Parity with going direct

Two mechanisms do not save against going direct. They keep a request through Context Mode as cheap as the same request sent direct.

Where you see the saving

The console's ledger prices each request twice: at list price for the usage the provider returned, and for what Claude Code going direct would have sent, counted with the provider's own token counter. The saving is the difference between the two. A request gets a dollar figure only when its count is within 5% of the provider's prompt count.

Thinking in Code, in depth

Thinking in Code gives the model two tools. The model writes a short program, the program reads the data, and only what the program returns enters the conversation.

context-mode-run-codeJavaScript in a Cloudflare Worker isolate, run by the gateway. Your agent never sees the call, so its permission rules and hooks do not apply to it. Egress is allow-listed, and every host the program calls is checked against your Cage rules when Cage is on. For public web pages and APIs.
context-mode-run-localThe gateway turns the call into your agent's own shell call: Bash in Claude Code, exec in Codex. The program runs on your machine, as a normal shell command. What it prints goes through the gateway to the model, and long output is archived. For local files and localhost.

Before and after

The Data Analysis benchmark asks about the npm registry documents for react and express. Below is one question, measured by hand with the script under it on 2026-09-29 (UTC). It is one query, not a benchmark figure.

What the fetch tool passes on
100,000 characters

Claude Code's fetch tool hands the first 100,000 characters of a page to a helper model: about 1.4% of the 7,011,585-byte react document (2,959 versions). Sizes change as new versions ship.

What the script returns
407 bytes

The version count, the dist-tags, and the first and latest stable release with their dates, computed over the whole document.

// Node 18 or later, or context-mode-run-code. Prints the bytes in, the bytes out and the answer.
const r = await fetch("https://registry.npmjs.org/react");
const t = await r.text(), d = JSON.parse(t);
const stable = Object.keys(d.versions).filter((v) => !v.includes("-"))
  .sort((a, b) => new Date(d.time[a]) - new Date(d.time[b]));
const latest = d["dist-tags"].latest;
const out = JSON.stringify({ versions: Object.keys(d.versions).length, dist_tags: d["dist-tags"],
  first_stable: { version: stable[0], date: d.time[stable[0]] },
  latest_stable: { version: latest, date: d.time[latest] } });
const bytes = (x) => new TextEncoder().encode(x).length;
console.log(bytes(t), "bytes in,", bytes(out), "bytes out");
console.log(out);

Without Thinking in Code, the agent has two choices. It can read the page through its fetch tool, which sees only the first part of the document. Or it can write its own download and script in the shell, where every command and its output stay in the conversation. With Thinking in Code, the program reads the whole document outside the conversation and returns only the answer.

How the model decides

The gateway adds the two tools to each request. Each tool description says when to use it and when not: run-code for a public page or API that has to be fetched and parsed, or for work that can be computed instead of read; run-local for files on your machine and localhost. Nothing forces the tool, with one exception: when steering expects a large output, it declines the shell call once and asks for run-local (see Steering). Otherwise the model picks the tool when the task fits, the same way it picks any other tool.

Why it is safe

Measured value

All Context Saving mechanisms were on in these runs. This is the task where Thinking in Code does most of the work.

Long-Context Data AnalysisSavedDirectContext ModePer 1,000 tasksHead-to-headCorrect answers
Claude Haiku 4.564%Haiku 4.5, likely 52 to 77%$0.175$0.062$11312 of 1211 of 12 in each arm
Claude Opus 5.510%Opus 5.5, likely 4 to 15%$0.223$0.201$2215 of 2020 of 20 in each arm

Business value. On Haiku 4.5 a data-analysis task cost $0.062 through Context Mode against $0.175 direct: $113 less per 1,000 tasks. On Opus 5.5 it was $22 less per 1,000 tasks. The Context Mode figures include the extra model calls the gateway makes when it answers run-code on its side.

Why the two models differ. In our earlier runs of this task (2026-09-28), Opus 5.5 going direct often wrote its own download and script, so less raw data reached its prompt and the gap is smaller. Haiku 4.5 going direct read more of the documents into its prompt. Thinking in Code pays most where a model would otherwise read the data.

The pattern is public. Anthropic's engineering post on code execution with MCP and Cloudflare's Code Mode post both describe a model writing code against its tools and reading back only the result. Context Mode runs it on the request path of the agents your team already uses, with Cage checks and a full archive. Product page: Thinking in Code.

Try it in 5 minutes

You run a big test suite and ask one question about it. The agent gets the answer. The full test log stays out of the conversation.

Setup

npx @context-mode/cli      # sign in and connect Claude Code, then restart Claude Code
git clone https://github.com/expressjs/express.git && cd express && npm install
claude

Prompt

Run the full test suite and tell me if anything fails.

What you will see

All tests pass — 1261 tests total.

⏺ context-mode · context 67 KB lighter than your raw transcript · kept 71 KB out of what the model read

We ran this on 2026-10-04 in real Claude Code sessions (claude -p 2.1.283, Claude Haiku 4.5) through the gateway, each on a new test account. Three runs. All three gave the right test count, or the right failing tests.

Why it matters. Test runs, builds and logs are the biggest things an agent reads in a day. The model gets the answer, and the context stays small for the rest of the session.

Compared with other approaches

Each approach below solves part of the problem. The table says what each one does and does not do, as we understand it.

ApproachWhat it doesWhat it does not do
Truncation and preview in the clientCuts a large output or saves it to a file and shows a preview.Output under the cut line enters whole and stays for the session. What was cut is not in the conversation, and the preview does not say which lines matter.
Summarising old turnsReplaces a long conversation with a written summary to free room.The summary is lossy: exact output, ids and numbers can go. The next request writes the new start to cache again.
RAG and indexingStores documents and returns the passages that match a query.It does not shrink tool output that is already in the conversation. The model still needs the right query, and a passage is not a computed answer.
Output filters on your machineShell wrappers or hooks shorten command output before the agent sees it.They run per machine and per client. Unless they keep a copy, what they cut is gone. They cannot move cache marks or fit requests to the window.
Provider prompt caching aloneReads a repeated prompt start at a lower price.It does not make the prompt smaller. It only pays when the cached text stays the same, so any edit of earlier text is a new cache write.
Context ModeActs on the request itself: first-sight folding with a full archive, code runs that return only the answer, cache-mark placement, and window fitting.It does not edit your files or your agent's settings. It rewrites cached text only at fixed steps on very long threads, one request per step.

Evaluation

Method

Results by category

Category and taskSavedDirectContext ModePer 1,000 tasksHead-to-headEvidence
Long-Context Data Analysis
Questions about the react and express npm registry documents, answered with a script64%Haiku 4.5, likely 52 to 77%$0.175$0.062$11312 of 12JSON · Pre-registration
Questions about the react and express npm registry documents, answered with a script10%Opus 5.5, likely 4 to 15%, later Gateway commit$0.223$0.201$2215 of 20JSON · Pre-registration
Test-Driven Debugging
npm test fix loop: run the tests, fix the code, run them again53%Haiku 4.5, likely 49 to 56%$0.069$0.032$3612 of 12JSON · Pre-registration
npm test fix loop, short22%Opus 5.5, likely 16 to 28%$0.087$0.068$1915 of 16JSON · Pre-registration
npm test fix loop, long21%Opus 5.5, likely 14 to 28%$0.085$0.067$1811 of 12JSON · Pre-registration
Agentic Coding (short, one-prompt coding tasks)
Code search (grep over 150 files)29%Opus 5.5, likely 23 to 34%$0.093$0.066$2716 of 16JSON · Pre-registration
One-line edit28%Opus 5.5, likely 14 to 42%$0.075$0.054$2111 of 12JSON · Pre-registration
Repo history question (git log, 320 commits)25%Opus 5.5, likely 16 to 34%$0.071$0.053$1815 of 16JSON · Pre-registration
Code search (grep)29%Haiku 4.5, likely 4 to 54%$0.031$0.022$918 of 20JSON · Pre-registration
One-line edit20%Haiku 4.5, likely 15 to 25%$0.025$0.020$58 of 8JSON · Pre-registration
Repo history question (git log)19%Haiku 4.5, likely 16 to 23%$0.025$0.020$512 of 12JSON · Pre-registration
Large Tool-Output Handling
A question about one large log22%Opus 5.5, likely 15 to 30%$0.091$0.071$2015 of 16JSON · Pre-registration
A question about one large JSON file21%Opus 5.5, likely 13 to 30%$0.086$0.068$1812 of 12JSON · Pre-registration
A question about one large JSON file16%Haiku 4.5, likely 9 to 23%$0.032$0.027$518 of 20JSON · Pre-registration
Multi-Agent Orchestration
Three parallel subagents, then one more15%Opus 5.5, likely 12 to 19%$0.481$0.408$7320 of 20JSON · Pre-registration

Cost at list price, paired runs, 2026-09-29, Claude Code 2.1.283. Cost per task is the mean cost. Per 1,000 tasks is the mean saving per task, times 1,000. Head-to-head counts the pairs where Context Mode cost less. We also ran the test-fix loop in Ruby to check the result is not tied to one toolchain: Opus 5.5 61% lower (likely 52 to 70%), Haiku 4.5 46% lower (likely 34 to 57%), 8 of 8 pairs won on each model.

Cost basis: provider usage at list price, not the client's estimate

Cost is list price applied to the token counts Anthropic returned for each call. The runs used a claude.ai login, which has no per-token invoice.

Through a gateway, Claude Code's own figure is wrong in two ways, so we do not use it for the Context Mode arm. It prices the WebFetch helper at the main model's price, although the call ran on Haiku 4.5. And it does not see the calls the gateway makes when it answers run-code on its side.

The check that the gateway record misses nothing: over the 224 valid Context Mode runs on gateway build 3bedf1ba, the gateway record holds at least every token Claude Code saw, in each class. In the 172 runs with no gateway-side call, the two are equal. On Data Analysis (Haiku 4.5, and an earlier Opus 5.5 run), in 32 of 32 runs the difference is exactly the gateway-side calls at list price. Pair 1 of that earlier Opus 5.5 Data Analysis run, call by call:

Context Mode arm, pair 1CallsAt list price
Calls Claude Code made, on Claude Opus 5.5 (Claude Code's own figure)3$0.156003
Calls the gateway made on its side, after it answered run-code2$0.050448
Context Mode arm, all calls5$0.206451
Direct arm, pair 1, Claude Code's own total$0.235333

Rows are rounded to the sixth decimal. The figures for this pair are in the published evidence file: "Large documents processed with a script", pair 1, $0.2353328 direct and $0.206451 through Context Mode. Claude Code's own figure is lower than the gateway record because it misses the gateway-side calls; we publish the gateway record.

Safety and limits

What it keeps

When the saving is small

What it does not do

Reproduce the numbers

Every result on this page comes from an evidence file on this site. Each file lists every pair with the cost per arm and the success check. The evidence index gives the method, the client version, the gateway version and what we removed (prompt text, session and tenant ids). This script recomputes a row from a file:

import json, statistics as st, sys
T = {7: 2.365, 11: 2.201, 15: 2.131, 19: 2.093}   # t at 97.5%, by degrees of freedom
path, task = sys.argv[1], sys.argv[2]            # e.g. final-haiku45.json W3
w = json.load(open(path))["workloads"]
w = w[task] if isinstance(w, dict) else next(x for x in w if x["id"] == task)
P = [p["plain"]["usd"] for p in w["pairs"]]
D = [p["gateway"]["usd"] - p["plain"]["usd"] for p in w["pairs"]]
n, mp, md = len(D), st.mean(P), st.mean(D)
h = T[n - 1] * st.stdev(D) / n ** 0.5
print(task, n, "pairs:", round(100 * md / mp, 1), "% (likely",
      round(100 * (md + h) / mp, 1), "to", round(100 * (md - h) / mp, 1), "%)",
      "won", sum(d < 0 for d in D))

Run on final-haiku45.json W3, it prints 12 pairs, −64.4%, likely −52.3 to −76.5%, won 12. On v4-opus55-W1-W2-W3-W4.json W2 it prints 20 pairs, −15.2%, likely −11.7 to −18.7%, won 20. Your own saving depends on your mix of work; the console measures it from your traffic.

FAQ for CTOs

Does a proxy break my prompt cache?

It is built not to. A result is changed only the first time it passes, and the same bytes go out on every later request. The gateway never removes your client's cache marks and never sends more than the provider allows. It rewrites cached text only at fixed steps on very long threads. The results are measured cost, so every cache write is in them.

Does quality drop?

Task success was equal in both arms on every published result, with a check per task (tests pass, the right answer). Test and build filters keep failures, errors and summaries as printed.

Is this truncation that loses data?

No. The full text is archived before the short view is made, and the short view names its archive id. The agent reads the text back with context-mode-search. Secrets and e-mail addresses are masked in the archive.

Did you pick the best benchmarks?

We publish tasks where the whole 95% interval is below zero. Each run's tasks, pair count and analysis were committed before its first pair, and no pair was dropped. Each result row links its pre-registration, with the commit and its date. The published rows ran on three gateway builds on 2026-09-29. When we fixed a defect, we ran the row again on the new build, pre-registered the same way, and the new run replaced the old one. Benchmarks lists the sources.

Aren't 8 to 20 pairs too few?

The design is paired, so each pair is its own control, and the interval is a t interval over the paired differences. Head-to-head counts are shown with each row: Multi-Agent won 20 of 20 pairs on Opus 5.5.

Is that your meter, not my bill?

The runs used a claude.ai login, so there is no per-token invoice to show. Both arms are list price applied to the token counts Anthropic returned. The direct arm is Claude Code's own total. The Context Mode arm is the gateway's per-call record, which holds at least every token Claude Code saw. See Cost basis.

Is this production data?

No. It is a controlled A/B with the real client (claude -p), not a replay. Your saving depends on your work; the console measures it from your own traffic.

Is a proxy in my path a risk?

Leaving takes one step: remove ANTHROPIC_BASE_URL from ~/.claude/settings.json; the CLI kept a backup. A change to the request shape must pass a real-client test through the gateway before it ships. See How do I leave?

What about privacy?

run-local runs on your machine; what it prints goes through the gateway to the model, and long output is archived. Each run-code network call is allow-listed and checked against Cage when Cage is on. Archives live on Cloudflare, one store per account. Prompts have no automatic expiry yet, and there is no one-click delete or export yet; ask us and we delete your data. See Data.

Won't the model provider ship this?

The code-execution pattern is public, and we credit it. The results on this page come from work on the request itself: Thinking in Code, first-sight folding with a full archive, cache-mark placement and page caching. A plugin inside the client cannot move cache marks or change the request the client sends. Some results depend on how Claude Code 2.1.283 builds its requests, so a client release can change them; each result names its client version.

Compared with other tools: see the landscape.