# Your agent pays for the same test log on every call

October 9, 2026 Saving 6 min read

A coding agent sends the whole conversation to the model on every call, so a test log it read once is paid for again and again. Context Mode folds that output the first time it comes back, keeps every byte in your archive, and in paired runs cut a 30-step Opus 5.5 session from $7.20 to $2.49.

## The problem: you pay for old output on every call

You ask your agent to run the tests. It reads the result, fixes one thing, and moves on. You think that test run is done. It is not.

A coding agent does not send the model one message at a time. It sends the whole conversation on every call: the system prompt, the tool schemas, every earlier turn and every tool result so far. A test run the agent read at step 3 is still in the request at step 30.

Four things set what a session costs:

- **Tool output stays.** A test run, a log or a JSON file enters the conversation and rides along on every later call. Claude Code saves Bash output over 30,000 characters to a file and shows a preview. Output under that line enters the prompt whole.
- **Re-reads.** Agents often read the same thing twice.
- **Cache rewrites.** The provider caches the start of the prompt. If something edits text inside that part, the next call writes it to cache again.
- **Start cost.** Tool schemas and the system prompt sit at the head of every request. In one recording (131 request bodies, 2026-08-14), a trivial 148,034-byte turn held 105,258 bytes of tool schemas and 27,951 bytes of system prompt. The user's own words were 234 bytes.

Cache prices set the rules. On Claude Opus 5.5 the list price per million tokens is $4 for fresh input, $8 for a 1-hour cache write, $0.20 for a cache read and $20 for output ([Anthropic prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching)). A cached token costs a twentieth of a fresh one. A token written to cache costs twice a fresh one. So a change that rewrites cached text can cost more than it saves. Context Mode is built around that fact.

## The fix: pay for the failing test, not the 599 that passed

Context Mode sits between your agent and the model. When a shell or search result comes back (Bash and Grep in Claude Code, exec and exec_command in Codex), it runs 15 filters on it, in a fixed order, before the model reads it.

Take the common case. Your agent runs `npm test`. 600 tests, 1 fails.

**What the test runner printed** *17.0 KB* **What the model reads** *1.2 KB*

The failing test, its assertion and the summary line stay exactly as printed. The 599 passing lines become one counted line.

The model sees this in place of the passing tests:

```
… 599 passing tests folded (599 lines)
```

A label after it names an archive id of the form `context-mode-toolcap-<hash>-<length>`. The same filter on a Ruby `minitest -v` run with 401 tests and 1 failure took 16.2 KB to 0.7 KB. Both sizes come from the test fixtures of our folding A/B (Claude Haiku 4.5, 2026-09-26).

The other filters work the same way:

- Progress redraws keep only their last frame, and colour codes go.
- Install and build progress folds into counted lines.
- A run of library stack frames becomes one line. Your own frames stay.
- Repeated lines are counted once.
- Search matches and build errors are grouped by file.

## What folding never drops

- **Nothing is dropped.** The full output goes to your archive before the short view is made. The agent reads it back with `context-mode-search` and the id, with an offset for later pages.
- **Failures stay.** Test and build filters keep failing tests, error lines and summaries as printed. A line no filter knows stays as printed.
- **Small output passes through.** Results under 1,024 characters are not touched. A fold is kept only if it saves at least 768 characters and 10% of the result.
- **File reads are never folded.** The model edits files from what Read returns, so Read and tool-server (MCP) results skip folding.
- **First sight only.** Context Mode changes a result only the first time it passes, and the change depends only on its text. Every later request carries the same bytes, so the provider keeps reading them from cache.

## When the output is data: read 7 MB, send 407 bytes

Folding helps when output is noise. Some output is data the agent must compute over. For that, Context Mode adds two tools:

- `context-mode-run-code` runs JavaScript in a sandboxed isolate, for public pages and APIs.
- `context-mode-run-local` turns into your agent's own shell call, for local files.

The model writes a short program, the program reads the data, and only what it returns enters the conversation. We call this Thinking in Code.

One hand-measured example (one query on 2026-09-29, not a benchmark figure). Question: how many versions of react are on npm?

- The npm registry document for react was 7,011,585 bytes, with 2,959 versions.
- Claude Code's fetch tool passes the first 100,000 characters of a page to a helper model: about 1.4% of that document.
- A 10-line script read the whole document and returned 407 bytes: the version count, the dist-tags, and the first and latest stable release with their dates.

The pattern is public: Anthropic's post on code execution with MCP and Cloudflare's Code Mode post both describe it. What we add is running it inside the agent you already use, with [Protect](https://context-mode.com/protect) checks on every host the script calls and a full archive of the output.

## Proof: paired benchmarks in dollars

Bytes are not dollars. The bill depends on cache reads, cache writes and how many turns the agent takes. So we run paired A/B tests: the same task with Claude Code going direct to Anthropic and through Context Mode, back to back, with the order switching each pair.

- Each row's plan was committed before its first pair.
- Cost is list price on the token counts Anthropic returned, call by call, including Context Mode's own extra calls.
- Client: Claude Code 2.1.283. Every row had equal task success in both arms.

| Workload (Opus 5.5 unless noted) | Pairs | Direct | Context Mode | Lower by | 95% range | Pairs cheaper |
| --- | --- | --- | --- | --- | --- | --- |
| npm test fix loop | 16 | $0.087 | $0.068 | 22% | 16 to 28% | 15 of 16 |
| One large log | 16 | $0.091 | $0.071 | 22% | 15 to 30% | 15 of 16 |
| One large JSON file | 12 | $0.086 | $0.068 | 21% | 13 to 30% | 12 of 12 |
| Data analysis with a script, Haiku 4.5 | 12 | $0.175 | $0.062 | 64% | 52 to 77% | 12 of 12 |
| Data analysis with a script | 20 | $0.223 | $0.201 | 10% | 4 to 15% | 15 of 20 |

Why the two data-analysis rows differ: in earlier runs, Opus 5.5 going direct often wrote its own download script, so less raw data reached its prompt. Haiku 4.5 read more of the documents. Thinking in Code pays most where a model would otherwise read the data.

## Long sessions: $7.20 to $2.49

Savings grow with session length, because every byte kept out would have been sent again on every later call. Our long-session test is one session of 30 steps on Opus 5.5 with the 1M window: read 16 source files, edit, test, record two decisions, then answer 10 memory questions. 4 pairs, run on 2026-10-03, with no compaction in any arm.

**Direct to Anthropic** *$7.20* **Through Context Mode** *$2.49*

Cost at list price per session, mean of 4 pairs. 65.5% lower (64.5% to 67.7% across pairs); all 4 pairs cost less.

| Per session (mean of 4 pairs) | Direct | Context Mode |
| --- | --- | --- |
| Cost at list price | $7.20 | $2.49 |
| Tokens sent per call, mean | 330,785 | 75,591 |
| Tokens sent per call, max | 438,707 | 126,750 |
| Task checks | 8 of 8 | 8 of 8 |

Two things to read with it:

- **Exact recall was lower in this run.** The questions forbade every tool, including the archive search. Context Mode had moved some early file reads to the archive, so it answered 5 to 6 of 10 questions exactly, against 10 of 10 direct. In a follow-up with the archive search allowed, it cost $2.97 against $7.24 (59.0% lower) and answered 10 of 10 exactly in three pairs and 8 of 10 in the fourth, where both misses held the right value with extra words.
- **The 65.5% belongs to all of Context Mode, not one feature.** Every feature was on in every run. No run isolates folding or Thinking in Code.

## What is not measured yet

- **Codex cost.** Tool-schema trimming now runs for Codex as well as Claude Code, and the console measures Codex cost per request. The Codex A/B has not run, so we quote no Codex saving.
- **Short work.** A short question with few tool calls and little output leaves little to save, so expect a small or no change there.
- **Your traffic.** These are controlled runs, not production data. The console measures what each request kept out on your own work.

## How to try it

```
npx @context-mode/cli          # sign in and connect Claude Code, then restart it
git clone https://github.com/expressjs/express.git && cd express && npm install
claude
```

Ask: "Run the full test suite and tell me if anything fails." `npm test` prints about 70 KB here, one line per test. In our runs (2026-10-04, Claude Haiku 4.5, three runs on new test accounts), the agent reported 1261 passing tests and the status line said `kept 71 KB out`. Your numbers will differ: a second run said 65 KB, and a run where the agent grepped the log first said 5 KB.

- Setup: [Quick start](https://context-mode.com/docs/quick-start).
- Every filter and rule: [How saving works](https://context-mode.com/docs/how-saving-works) and [Context saving](https://context-mode.com/context-saving).
- Scripts over data: [Thinking in Code](https://context-mode.com/thinking-in-code).
- Full tables and method: [Benchmarks](https://context-mode.com/docs/benchmarks).
- Plans: [Plans and limits](https://context-mode.com/docs/plans-and-limits). Free gives 1,000 requests once, with every feature. Your own Anthropic or OpenAI account pays the model, and we add no fee on tokens.

To leave, remove `ANTHROPIC_BASE_URL` from `~/.claude/settings.json`.

[Start free](https://context-mode.com/docs/quick-start)[All posts](https://context-mode.com/blog)
