Your agent pays for the same test log on every call

Saving6 min read

A coding agent sends the whole conversation to the model on every call, so a test log it read once is paid for again and again. Context Mode folds that output the first time it comes back, keeps every byte in your archive, and in paired runs cut a 30-step Opus 5.5 session from $7.20 to $2.49.

The problem: you pay for old output on every call

You ask your agent to run the tests. It reads the result, fixes one thing, and moves on. You think that test run is done. It is not.

A coding agent does not send the model one message at a time. It sends the whole conversation on every call: the system prompt, the tool schemas, every earlier turn and every tool result so far. A test run the agent read at step 3 is still in the request at step 30.

Four things set what a session costs:

Cache prices set the rules. On Claude Opus 5.5 the list price per million tokens is $4 for fresh input, $8 for a 1-hour cache write, $0.20 for a cache read and $20 for output (Anthropic prompt caching). A cached token costs a twentieth of a fresh one. A token written to cache costs twice a fresh one. So a change that rewrites cached text can cost more than it saves. Context Mode is built around that fact.

The fix: pay for the failing test, not the 599 that passed

Context Mode sits between your agent and the model. When a shell or search result comes back (Bash and Grep in Claude Code, exec and exec_command in Codex), it runs 15 filters on it, in a fixed order, before the model reads it.

Take the common case. Your agent runs npm test. 600 tests, 1 fails.

What the test runner printed17.0 KB
What the model reads1.2 KB
The failing test, its assertion and the summary line stay exactly as printed. The 599 passing lines become one counted line.

The model sees this in place of the passing tests:

… 599 passing tests folded (599 lines)

A label after it names an archive id of the form context-mode-toolcap-<hash>-<length>. The same filter on a Ruby minitest -v run with 401 tests and 1 failure took 16.2 KB to 0.7 KB. Both sizes come from the test fixtures of our folding A/B (Claude Haiku 4.5, 2026-09-26).

The other filters work the same way:

What folding never drops

When the output is data: read 7 MB, send 407 bytes

Folding helps when output is noise. Some output is data the agent must compute over. For that, Context Mode adds two tools:

The model writes a short program, the program reads the data, and only what it returns enters the conversation. We call this Thinking in Code.

One hand-measured example (one query on 2026-09-29, not a benchmark figure). Question: how many versions of react are on npm?

The pattern is public: Anthropic's post on code execution with MCP and Cloudflare's Code Mode post both describe it. What we add is running it inside the agent you already use, with Protect checks on every host the script calls and a full archive of the output.

Proof: paired benchmarks in dollars

Bytes are not dollars. The bill depends on cache reads, cache writes and how many turns the agent takes. So we run paired A/B tests: the same task with Claude Code going direct to Anthropic and through Context Mode, back to back, with the order switching each pair.

Workload (Opus 5.5 unless noted)PairsDirectContext ModeLower by95% rangePairs cheaper
npm test fix loop16$0.087$0.06822%16 to 28%15 of 16
One large log16$0.091$0.07122%15 to 30%15 of 16
One large JSON file12$0.086$0.06821%13 to 30%12 of 12
Data analysis with a script, Haiku 4.512$0.175$0.06264%52 to 77%12 of 12
Data analysis with a script20$0.223$0.20110%4 to 15%15 of 20

Why the two data-analysis rows differ: in earlier runs, Opus 5.5 going direct often wrote its own download script, so less raw data reached its prompt. Haiku 4.5 read more of the documents. Thinking in Code pays most where a model would otherwise read the data.

Long sessions: $7.20 to $2.49

Savings grow with session length, because every byte kept out would have been sent again on every later call. Our long-session test is one session of 30 steps on Opus 5.5 with the 1M window: read 16 source files, edit, test, record two decisions, then answer 10 memory questions. 4 pairs, run on 2026-10-03, with no compaction in any arm.

Direct to Anthropic$7.20
Through Context Mode$2.49
Cost at list price per session, mean of 4 pairs. 65.5% lower (64.5% to 67.7% across pairs); all 4 pairs cost less.
Per session (mean of 4 pairs)DirectContext Mode
Cost at list price$7.20$2.49
Tokens sent per call, mean330,78575,591
Tokens sent per call, max438,707126,750
Task checks8 of 88 of 8

Two things to read with it:

What is not measured yet

How to try it

npx @context-mode/cli          # sign in and connect Claude Code, then restart it
git clone https://github.com/expressjs/express.git && cd express && npm install
claude

Ask: "Run the full test suite and tell me if anything fails." npm test prints about 70 KB here, one line per test. In our runs (2026-10-04, Claude Haiku 4.5, three runs on new test accounts), the agent reported 1261 passing tests and the status line said kept 71 KB out. Your numbers will differ: a second run said 65 KB, and a run where the agent grepped the log first said 5 KB.

To leave, remove ANTHROPIC_BASE_URL from ~/.claude/settings.json.