Claude Code cost benchmarks

Through Context Mode Gateway, Claude Code on Claude Opus 5.5 cost 65.5% less on a long agentic coding session, which went from $7.20 to $2.49, 4 of 4 pairs cheaper. On the 10 shorter tasks shown here, 155 of 168 paired runs cost less, and 168 of 168 tasks succeeded in both arms. Each pair ran the same task with Claude Code direct to Anthropic and through the gateway, back to back, with the plan committed before the first pair.

Key facts

What it is
Paired A/B cost tests of Claude Code direct and through Context Mode Gateway, each run with a pre-registered plan.
Who it is for
CTOs and engineers who want to check our cost claims before they connect an agent.
Clients
Claude Code 2.1.283, on Claude Opus 5.5 and Claude Haiku 4.5. Codex is not measured yet.
Measured result
Opus 5.5, 2026-10-03: 65.5% lower on a long session, $7.20 to $2.49, 4 of 4 pairs cheaper. On 2026-09-29, 155 of 168 shorter pairs cost less. All published pairs.
Limits
This page shows the workloads where cost was lower; the pre-registrations mark where workloads were taken off. Cost is list price on the token counts Anthropic returned: cash on API billing, list-price value on a Claude plan.

Same task success, lower cost. Task checks passed equally in both arms on every row below. The long-session rows also ask 10 memory questions; each row gives both arms' scores.

Claude Opus 5.5

Cost per task is the mean cost at list price. "Likely" gives the 95% range. Saved per 1,000 tasks is the mean saving per task, times 1,000.

Category and taskSavedDirectContext ModePer 1,000 tasksHead-to-headEvidence
Agentic CodingEveryday agent steps: search the code, ask the repo history, make an edit.
Code search (grep over 150 files)29%Opus 5.5, likely 23 to 34%$0.093$0.066$2716 of 16JSON · Pre-registration
One-line edit28%Opus 5.5, likely 14 to 42%$0.075$0.054$2111 of 12JSON · Pre-registration
Repo history question (git log, 320 commits)25%Opus 5.5, likely 16 to 34%$0.071$0.053$1815 of 16JSON · Pre-registration
Large Tool-Output HandlingA command prints far more than the model needs.
One large log22%Opus 5.5, likely 15 to 30%$0.091$0.071$2015 of 16JSON · Pre-registration
One large JSON file21%Opus 5.5, likely 13 to 30%$0.086$0.068$1812 of 12JSON · Pre-registration
Test-Driven DebuggingRun the tests, fix the code, run them again.
npm test fix loop, short22%Opus 5.5, likely 16 to 28%$0.087$0.068$1915 of 16JSON · Pre-registration
npm test fix loop, long21%Opus 5.5, likely 14 to 28%$0.085$0.067$1811 of 12JSON · Pre-registration
Multi-Agent OrchestrationA lead agent splits the work across parallel subagents.
Three parallel subagents, then one more15%Opus 5.5, likely 12 to 19%$0.481$0.408$7320 of 20JSON · Pre-registration
Long Agentic Coding SessionOne session of 30 steps: read 16 files, edit, test, then 10 memory questions. Run on the 1M window and on a 200K window.
Long agentic coding session, Opus 5.5, 1M window65.5%Opus 5.5, lowest to highest pair 64.5 to 67.7%$7.20$2.49$4,7154 of 4Archive search off in this run; the run with it on is in the details. DetailsJSON · Evidence index
Long agentic coding session, Opus 5.5, 200K window, Recall on37.6%Opus 5.5, 3 pairs, 30.7 to 44.3%$4.86$3.04$1,8273 of 3Memory answers 9 of 10, direct 4 of 10JSON · Evidence index
Long-Context Data AnalysisThe agent writes a script to process large documents.
Large documents processed with a script10%Opus 5.5, likely 4 to 15%, later Gateway commit$0.223$0.201$2215 of 20JSON · Pre-registration
Large documents processed with a script9%Opus 5.5, likely 5 to 14%, earlier Gateway commit$0.226$0.205$2117 of 20JSON · Pre-registration

Three parallel subagents: the 15% and 20 of 20 come from the gateway's record, and Claude Code's own cost figures for both arms also give 15% and 20 of 20. The gateway did not save how long each cache write lived, so from the saved tokens alone a reader can confirm 5 to 23% lower, with 15 to 20 of 20 pairs lower.

Long agentic coding session: what else we measured

4 pairs on 2026-10-03, Claude Code 2.1.283 on Claude Opus 5.5 with its real 1M window, Recall on in the Context Mode arm. Each pair ran the same 30 steps in both arms. Every number links to the file it comes from.

Long session on a 200K window: savings and quality

3 pairs on 2026-10-04: the same 30 steps as above, with Claude Code 2.1.283 on Claude Opus 5.5 and its real 200K window, and the gateway account's Context limit at 100K. Recall on in the Context Mode arm. The rules were written before the first pair, and all 3 of 3 pairs are valid. What that means.

All published Opus 5.5 short-task pairs: 155 of 168 cost less

Every Opus 5.5 pair in the three published evidence files, added up: 168 paired runs over 10 tasks, $28.28 direct and $23.16 through Context Mode at list price. 155 of 168 pairs cost less, and 168 of 168 tasks succeeded in both arms. It is a sum over pairs, not an average of rows. These are the 10 tasks where cost was lower: workloads without a lower-cost result are not in the published files. On API billing it is cash; on a Claude plan it is list-price value, not cash.

TaskPairsCost lessEvidence
Code search (grep over 150 files)1616 of 16JSON · Pre-registration
One-line edit1211 of 12JSON · Pre-registration
Repo history question (git log, 320 commits)1615 of 16JSON · Pre-registration
One large log1615 of 16JSON · Pre-registration
One large JSON file1212 of 12JSON · Pre-registration
npm test fix loop, short1615 of 16JSON · Pre-registration
npm test fix loop, long1211 of 12JSON · Pre-registration
Three parallel subagents, then one more2020 of 20JSON · Pre-registration
Large documents processed with a script, two runs of 204032 of 40JSON · Pre-registration · JSON · Pre-registration
Test fix loop (minitest; the note under How we measured)88 of 8JSON · Pre-registration
All together168155 of 168

Claude Haiku 4.5

The same tasks on a smaller, cheaper model.

Category and taskSavedDirectContext ModePer 1,000 tasksHead-to-headEvidence
Long-Context Data Analysis
Large documents processed with a script64%Haiku 4.5, likely 52 to 77%$0.175$0.062$11312 of 12JSON · Pre-registration
Test-Driven Debugging
npm test fix loop53%Haiku 4.5, likely 49 to 56%$0.069$0.032$3612 of 12JSON · Pre-registration
Agentic Coding
Code search (grep)29%Haiku 4.5, likely 4 to 54%$0.031$0.022$918 of 20JSON · Pre-registration
One-line edit20%Haiku 4.5, likely 15 to 25%$0.025$0.020$58 of 8JSON · Pre-registration
Repo history question (git log)19%Haiku 4.5, likely 16 to 23%$0.025$0.020$512 of 12JSON · Pre-registration
Large Tool-Output Handling
One large JSON file16%Haiku 4.5, likely 9 to 23%$0.032$0.027$518 of 20JSON · Pre-registration

Task success was equal in both arms on every Haiku 4.5 row. On data analysis, 11 of 12 answers were correct in each arm.

How we measured

We also ran the test-fix loop in a second language (Ruby) to check that the result is not tied to one toolchain: Opus 5.5 61% lower cost (likely 52 to 70%), Haiku 4.5 46% lower (likely 34 to 57%), 8 pairs each, 8 of 8 won on each model. We quote no figure for other models or for Codex. Your own saving depends on your mix of work; the console measures it from your traffic.

Verify it yourself

Every evidence file holds the token counts of each call and the list price we used, so you can recompute every dollar yourself. A checking script is available on request.

Sources

Every number on this page comes from an evidence file published on this site: each pair, its cost per arm, and the analysis. The files are copies of the committed files in the gateway repository, with private fields removed.

Launch benchmark, 2026-09-29The method, the evidence files and what was removed from them. Evidence index (JSON)
Long agentic coding session, 2026-10-03Every JSON file of the run and of its amendment, as committed, with the test tenant ids withheld. Evidence index (JSON), result, amendment 4 result.
Recall evidence, 2026-10-03 and 2026-10-04The 200K-window runs, the offline quality replays, the live checks and the read-speed figures, as committed, with tenant ids withheld. Evidence index (JSON), 200K result.
Pre-registrationsOne file per run, committed before its first pair: the tasks, the pairs and the analysis. Each shows its commit, the commit date and the run's start time. Opus 5.5 short tasks, Opus 5.5 test loops, Opus 5.5 multi-agent and data analysis, Haiku 4.5, long session, long session, amendment 4, 200K run 1, 200K run 2.
Evidence registerEvery cost claim with its status, including earlier runs this page replaces.

Full table of the rows where cost was lower, with each row's pairs, range and verdict: launch benchmark, full table.

FAQ

Does a gateway really lower Claude Code cost?

Yes: 65.5% lower on a long Opus 5.5 session, and 155 of 168 shorter Opus 5.5 pairs cost less. Your saving depends on your work, and the console measures it from your own traffic.

Was the plan fixed before the runs?

Yes. Each run's plan was committed before its first pair. See Sources.

Is this my bill?

It is list price applied to the token counts Anthropic returned for each call. On API billing it is cash; on a Claude plan it is list-price value, not cash.

Did you measure Codex?

Not yet. Every figure here is for Claude Code.