# How to reduce Claude Code costs

To cut Claude Code costs, send less tool output, fewer tool schemas and fewer cache rewrites on each call, because Claude Code sends the whole conversation again on every call. Do the free steps first, then route Claude Code through [Context Mode Gateway](https://context-mode.com/): a long Claude Code session on Claude Opus 5.5 cost [65.5% less](https://context-mode.com/benchmarks/long-session/ab.json), [$7.20](https://context-mode.com/benchmarks/long-session/ab.json) down to [$2.49](https://context-mode.com/benchmarks/long-session/ab.json).

**What to do:**

1. Do the free steps from Anthropic's cost guide: see [what you can change yourself](#yourself).
2. Connect Claude Code or Codex to [Context Mode Gateway](https://context-mode.com/) with one command: `npx @context-mode/cli`. See the [quick start](https://context-mode.com/docs/quick-start).
3. Watch your own tokens, cost and saving per request in the [console](https://console.context-mode.com).

## Key facts

What it is

A guide to where Claude Code spends tokens, what you can change yourself, and what Context Mode Gateway changes on the request path.

Who it is for

Developers and CTOs who find Claude Code too expensive, on API billing or on a Claude plan.

Clients

Claude Code. Context Mode Gateway also works with Codex, but Codex cost is not measured yet.

Measured result

Claude Code on Claude Opus 5.5 cost [65.5% less](https://context-mode.com/benchmarks/long-session/ab.json) on a long agentic coding session, which went from [$7.20](https://context-mode.com/benchmarks/long-session/ab.json) to [$2.49](https://context-mode.com/benchmarks/long-session/ab.json), 4 of 4 pairs cheaper, run on 2026-10-03. Over 168 shorter paired runs on 2026-09-29, 155 cost less and 168 of 168 tasks passed in both arms. [Method and evidence](https://context-mode.com/docs/benchmarks#long-session).

Limits

We measured cost at list price, not plan usage limits. A short question with few tool calls leaves little to save.

## Where Claude Code spends tokens

A coding agent sends the whole conversation to the model on every call. Four things decide the cost:

- **Tool output stays.** A test run, a log or a JSON file the agent reads enters the conversation and is sent again on every later call of the session.
- **Re-reads.** For WebFetch, Claude Code hands each fetched page to a small helper call with no cache mark, so each read of the page is charged as fresh input.
- **Cache rewrites.** The provider caches the start of the prompt. If a tool changes text inside that cached part, the next call writes it to cache again. See [Prompt caching](https://context-mode.com/docs/prompt-caching).
- **Start cost.** Tool schemas and the system prompt sit at the start of every request, and each new session writes part of that start to cache again.

Anthropic's own cost page puts the average at about $13 per developer per active day and $150 to $250 per developer per month across enterprise deployments ([Manage costs effectively](https://code.claude.com/docs/en/costs), read 2026-10-01).

## What you can change yourself

Anthropic's cost page lists steps that need no extra tool. Among them: manage context early, choose the right model, cut MCP server overhead, move instructions from CLAUDE.md into skills, adjust extended thinking, send verbose work to subagents, and write specific prompts ([Manage costs effectively](https://code.claude.com/docs/en/costs), read 2026-10-01). Do these first. They cost nothing.

## What Context Mode Gateway changes

The gateway sits between Claude Code and the model API. It changes what each request carries, and does not change your files or your agent's settings.

[Context Saving](https://context-mode.com/context-saving) Test, build, install and log output folded the first time the model sees it. The failing test stays; passing tests become one line. The full output is archived.

[Thinking in Code](https://context-mode.com/thinking-in-code) The agent writes a short script, the gateway runs it, and only what it prints reaches the model.

Tool schemas Long tool descriptions are trimmed, and the model can ask for the full definition when it needs it.

Cache marks The fixed part of the prompt start gets its own cache mark, and fetched pages are marked for a 5-minute cache. See [Prompt caching](https://context-mode.com/docs/prompt-caching).

[Context limit](https://context-mode.com/docs/context-limit) How much history a thread keeps: 200K, 500K or 1M. Older parts move to an archive the agent can read back.

The gateway adds its own tools, memory and skills, and those cost tokens. The results below count both sides. Every change is listed in [Context Saving](https://context-mode.com/docs/how-saving-works).

## What it saved

Paired runs against Claude Code going direct, on Claude Opus 5.5, 2026-09-29. Ranges are 95% intervals. Task success was equal in both arms on every row.

| Category | Task | Pairs | Gateway vs direct |
| --- | --- | --- | --- |
| [Agentic Coding](https://context-mode.com/docs/benchmarks#agentic-coding) | Code search ( `grep` over 150 files) | 16 | Opus 5.5: 29% lower cost (23 to 34%) |
| [Agentic Coding](https://context-mode.com/docs/benchmarks#agentic-coding) | One-line edit | 12 | Opus 5.5: 28% lower cost (14 to 42%) |
| [Agentic Coding](https://context-mode.com/docs/benchmarks#agentic-coding) | Repo history question ( `git log` , 320 commits) | 16 | Opus 5.5: 25% lower cost (16 to 34%) |
| [Large Tool-Output Handling](https://context-mode.com/docs/benchmarks#tool-output) | One large log | 16 | Opus 5.5: 22% lower cost (15 to 30%) |
| [Test-Driven Debugging](https://context-mode.com/docs/benchmarks#test-debugging) | `npm test` fix loop, short | 16 | Opus 5.5: 22% lower cost (16 to 28%) |
| [Multi-Agent Orchestration](https://context-mode.com/docs/benchmarks#multi-agent) | Three parallel subagents, then one more | 20 | Opus 5.5: 15% lower cost (12 to 19%) |

A reply-style block that asks for short answers was on in every Context Mode run. Cost is list price applied to the token counts Anthropic returned. [Every result for both models, with its method and evidence](https://context-mode.com/docs/benchmarks).

## On a Claude plan

We measured cost at list price, not plan usage limits. On API billing the saving is cash. On a Claude plan it is list-price value, not cash, and we make no claim about how long your plan's limits last.

## FAQ

### How much does Claude Code cost per developer?

Anthropic's cost page says about $13 per developer per active day, and $150 to $250 per developer per month across enterprise deployments ([source](https://code.claude.com/docs/en/costs), read 2026-10-01).

### What does Claude Code spend most tokens on?

Tool output that is sent again on every later turn, tool schemas, and cache writes at the start of each request. See [where an agent's tokens go](https://context-mode.com/docs/how-saving-works#problem).

### Will Context Mode lower my bill?

It depends on your work. On Claude Opus 5.5, cost was 25 to 29% lower on agentic coding and 22% lower on test-driven debugging. The [console](https://console.context-mode.com) measures your own saving from your own traffic.

### Do I have to change my agent or model?

No. One command sets the base URL, and your Claude account still pays the model. See [Quick start](https://context-mode.com/docs/quick-start).

### Does it help with Claude plan usage limits?

We measured cost at list price, not plan limits. On a Claude plan the saving is list-price value, not cash.

[Connect your agent](https://context-mode.com/docs/quick-start) [See the benchmarks](https://context-mode.com/docs/benchmarks) [How saving works](https://context-mode.com/docs/how-saving-works)
