Claude Code cost benchmarks
Through Context Mode Gateway, Claude Code on Claude Opus 5.5 cost 65.5% less on a long agentic coding session, which went from $7.20 to $2.49, 4 of 4 pairs cheaper. On the 10 shorter tasks shown here, 155 of 168 paired runs cost less, and 168 of 168 tasks succeeded in both arms. Each pair ran the same task with Claude Code direct to Anthropic and through the gateway, back to back, with the plan committed before the first pair.
Key facts
- What it is
- Paired A/B cost tests of Claude Code direct and through Context Mode Gateway, each run with a pre-registered plan.
- Who it is for
- CTOs and engineers who want to check our cost claims before they connect an agent.
- Clients
- Claude Code 2.1.283, on Claude Opus 5.5 and Claude Haiku 4.5. Codex is not measured yet.
- Measured result
- Opus 5.5, 2026-10-03: 65.5% lower on a long session, $7.20 to $2.49, 4 of 4 pairs cheaper. On 2026-09-29, 155 of 168 shorter pairs cost less. All published pairs.
- Limits
- This page shows the workloads where cost was lower; the pre-registrations mark where workloads were taken off. Cost is list price on the token counts Anthropic returned: cash on API billing, list-price value on a Claude plan.
Same task success, lower cost. Task checks passed equally in both arms on every row below. The long-session rows also ask 10 memory questions; each row gives both arms' scores.
Claude Opus 5.5
Cost per task is the mean cost at list price. "Likely" gives the 95% range. Saved per 1,000 tasks is the mean saving per task, times 1,000.
| Category and task | Saved | Direct | Context Mode | Per 1,000 tasks | Head-to-head | Evidence |
|---|---|---|---|---|---|---|
| Agentic CodingEveryday agent steps: search the code, ask the repo history, make an edit. | ||||||
Code search (grep over 150 files) | 29%Opus 5.5, likely 23 to 34% | $0.093 | $0.066 | $27 | 16 of 16 | JSON · Pre-registration |
| One-line edit | 28%Opus 5.5, likely 14 to 42% | $0.075 | $0.054 | $21 | 11 of 12 | JSON · Pre-registration |
Repo history question (git log, 320 commits) | 25%Opus 5.5, likely 16 to 34% | $0.071 | $0.053 | $18 | 15 of 16 | JSON · Pre-registration |
| Large Tool-Output HandlingA command prints far more than the model needs. | ||||||
| One large log | 22%Opus 5.5, likely 15 to 30% | $0.091 | $0.071 | $20 | 15 of 16 | JSON · Pre-registration |
| One large JSON file | 21%Opus 5.5, likely 13 to 30% | $0.086 | $0.068 | $18 | 12 of 12 | JSON · Pre-registration |
| Test-Driven DebuggingRun the tests, fix the code, run them again. | ||||||
npm test fix loop, short | 22%Opus 5.5, likely 16 to 28% | $0.087 | $0.068 | $19 | 15 of 16 | JSON · Pre-registration |
npm test fix loop, long | 21%Opus 5.5, likely 14 to 28% | $0.085 | $0.067 | $18 | 11 of 12 | JSON · Pre-registration |
| Multi-Agent OrchestrationA lead agent splits the work across parallel subagents. | ||||||
| Three parallel subagents, then one more | 15%Opus 5.5, likely 12 to 19% | $0.481 | $0.408 | $73 | 20 of 20 | JSON · Pre-registration |
| Long Agentic Coding SessionOne session of 30 steps: read 16 files, edit, test, then 10 memory questions. Run on the 1M window and on a 200K window. | ||||||
| Long agentic coding session, Opus 5.5, 1M window | 65.5%Opus 5.5, lowest to highest pair 64.5 to 67.7% | $7.20 | $2.49 | $4,715 | 4 of 4Archive search off in this run; the run with it on is in the details. Details | JSON · Evidence index |
| Long agentic coding session, Opus 5.5, 200K window, Recall on | 37.6%Opus 5.5, 3 pairs, 30.7 to 44.3% | $4.86 | $3.04 | $1,827 | 3 of 3Memory answers 9 of 10, direct 4 of 10 | JSON · Evidence index |
| Long-Context Data AnalysisThe agent writes a script to process large documents. | ||||||
| Large documents processed with a script | 10%Opus 5.5, likely 4 to 15%, later Gateway commit | $0.223 | $0.201 | $22 | 15 of 20 | JSON · Pre-registration |
| Large documents processed with a script | 9%Opus 5.5, likely 5 to 14%, earlier Gateway commit | $0.226 | $0.205 | $21 | 17 of 20 | JSON · Pre-registration |
Three parallel subagents: the 15% and 20 of 20 come from the gateway's record, and Claude Code's own cost figures for both arms also give 15% and 20 of 20. The gateway did not save how long each cache write lived, so from the saved tokens alone a reader can confirm 5 to 23% lower, with 15 to 20 of 20 pairs lower.
Long agentic coding session: what else we measured
4 pairs on 2026-10-03, Claude Code 2.1.283 on Claude Opus 5.5 with its real 1M window, Recall on in the Context Mode arm. Each pair ran the same 30 steps in both arms. Every number links to the file it comes from.
- Task checks. 8 of 8 in both arms, in every pair. No compaction and no upstream 4xx error in any of the 8 runs.
- Memory questions. In this run the agent could not search its archive. The run below lets it, the way the product works.
- Tokens sent per call. 330,785 direct, 75,591 through Context Mode (mean of the 4 pairs).
- The same session, archive search allowed. Context Mode cost $2.97 against $7.24 direct, 59.0% lower. Memory answers through Context Mode rose to 10, 10, 10 and 8 of 10, against 10 of 10 direct. What that means.
Long session on a 200K window: savings and quality
3 pairs on 2026-10-04: the same 30 steps as above, with Claude Code 2.1.283 on Claude Opus 5.5 and its real 200K window, and the gateway account's Context limit at 100K. Recall on in the Context Mode arm. The rules were written before the first pair, and all 3 of 3 pairs are valid. What that means.
- Savings. $3.04 per session through Context Mode against $4.86 direct: $1.83 saved, 37.6% lower (30.7 to 44.3%). Cheaper in 3 of 3 pairs.
- Quality. Exact answers to the memory questions: 9 of 10 through Context Mode (the value was in the answer in 10 of 10), 4 of 10 direct. Direct compacted 3 times; through Context Mode the session never compacted. Task checks 8 of 8 in both arms.
All published Opus 5.5 short-task pairs: 155 of 168 cost less
Every Opus 5.5 pair in the three published evidence files, added up: 168 paired runs over 10 tasks, $28.28 direct and $23.16 through Context Mode at list price. 155 of 168 pairs cost less, and 168 of 168 tasks succeeded in both arms. It is a sum over pairs, not an average of rows. These are the 10 tasks where cost was lower: workloads without a lower-cost result are not in the published files. On API billing it is cash; on a Claude plan it is list-price value, not cash.
| Task | Pairs | Cost less | Evidence |
|---|---|---|---|
Code search (grep over 150 files) | 16 | 16 of 16 | JSON · Pre-registration |
| One-line edit | 12 | 11 of 12 | JSON · Pre-registration |
Repo history question (git log, 320 commits) | 16 | 15 of 16 | JSON · Pre-registration |
| One large log | 16 | 15 of 16 | JSON · Pre-registration |
| One large JSON file | 12 | 12 of 12 | JSON · Pre-registration |
npm test fix loop, short | 16 | 15 of 16 | JSON · Pre-registration |
npm test fix loop, long | 12 | 11 of 12 | JSON · Pre-registration |
| Three parallel subagents, then one more | 20 | 20 of 20 | JSON · Pre-registration |
| Large documents processed with a script, two runs of 20 | 40 | 32 of 40 | JSON · Pre-registration · JSON · Pre-registration |
| Test fix loop (minitest; the note under How we measured) | 8 | 8 of 8 | JSON · Pre-registration |
| All together | 168 | 155 of 168 |
Claude Haiku 4.5
The same tasks on a smaller, cheaper model.
| Category and task | Saved | Direct | Context Mode | Per 1,000 tasks | Head-to-head | Evidence |
|---|---|---|---|---|---|---|
| Long-Context Data Analysis | ||||||
| Large documents processed with a script | 64%Haiku 4.5, likely 52 to 77% | $0.175 | $0.062 | $113 | 12 of 12 | JSON · Pre-registration |
| Test-Driven Debugging | ||||||
npm test fix loop | 53%Haiku 4.5, likely 49 to 56% | $0.069 | $0.032 | $36 | 12 of 12 | JSON · Pre-registration |
| Agentic Coding | ||||||
Code search (grep) | 29%Haiku 4.5, likely 4 to 54% | $0.031 | $0.022 | $9 | 18 of 20 | JSON · Pre-registration |
| One-line edit | 20%Haiku 4.5, likely 15 to 25% | $0.025 | $0.020 | $5 | 8 of 8 | JSON · Pre-registration |
Repo history question (git log) | 19%Haiku 4.5, likely 16 to 23% | $0.025 | $0.020 | $5 | 12 of 12 | JSON · Pre-registration |
| Large Tool-Output Handling | ||||||
| One large JSON file | 16%Haiku 4.5, likely 9 to 23% | $0.032 | $0.027 | $5 | 18 of 20 | JSON · Pre-registration |
Task success was equal in both arms on every Haiku 4.5 row. On data analysis, 11 of 12 answers were correct in each arm.
How we measured
- Paired A/B. Each pair runs the same task twice, back to back: once direct to Anthropic, once through Context Mode. The order changes every pair.
- The real client. Claude Code 2.1.283 (
claude -p) in both arms, with a fresh config and no plugin for every run. - The cost. List price applied to the token counts Anthropic returned for each call. The runs used a claude.ai login, which has no per-token invoice. Direct: Claude Code's own total. Context Mode: the gateway's per-call record, hidden calls included. See Cost basis.
- Written down first. For every row, the tasks, the number of pairs and the analysis were committed before the first pair ran. Within each row shown, no pair was dropped.
- Which rows. This page shows the rows where cost was lower. It is not every workload we ran: a pre-registration marked "edited" shows where workloads were taken off.
- Gateway settings. Each workload ran on its own new account, at the default settings, and no run changed them: memory on, skills on with none imported, and Cage off.
- Pairs. 8 to 20 per row, shown in the head-to-head column. One warm-up pair per task ran first and does not count.
- What counts as a win. A row counts only when the whole 95% range is below zero. Head-to-head counts the pairs where Context Mode cost less.
- When. 2026-09-29, on Claude Opus 5.5 and Claude Haiku 4.5.
We also ran the test-fix loop in a second language (Ruby) to check that the result is not tied to one toolchain: Opus 5.5 61% lower cost (likely 52 to 70%), Haiku 4.5 46% lower (likely 34 to 57%), 8 pairs each, 8 of 8 won on each model. We quote no figure for other models or for Codex. Your own saving depends on your mix of work; the console measures it from your traffic.
Verify it yourself
Every evidence file holds the token counts of each call and the list price we used, so you can recompute every dollar yourself. A checking script is available on request.
Sources
Every number on this page comes from an evidence file published on this site: each pair, its cost per arm, and the analysis. The files are copies of the committed files in the gateway repository, with private fields removed.
Full table of the rows where cost was lower, with each row's pairs, range and verdict: launch benchmark, full table.
FAQ
Does a gateway really lower Claude Code cost?
Yes: 65.5% lower on a long Opus 5.5 session, and 155 of 168 shorter Opus 5.5 pairs cost less. Your saving depends on your work, and the console measures it from your own traffic.
Was the plan fixed before the runs?
Yes. Each run's plan was committed before its first pair. See Sources.
Is this my bill?
It is list price applied to the token counts Anthropic returned for each call. On API billing it is cash; on a Claude plan it is list-price value, not cash.
Did you measure Codex?
Not yet. Every figure here is for Claude Code.