# Context Mode launch benchmark: the full table

The workloads measured on 2026-09-29 where cost was lower, each with a published evidence file. This is not every workload we ran: the pre-registrations marked "edited" show where workloads were taken off. Each row is a paired A/B test: the same task, run with Claude Code direct to Anthropic and through Context Mode, back to back.

## Method

- **Client.** Claude Code 2.1.283 (`claude -p`) in both arms, with a fresh config and no plugin for every run.
- **Pairs.** 8 to 20 per row, fixed and pre-registered before the first pair (each row links its pre-registration). One warm-up pair per workload ran first and does not count. The arm order changes every pair.
- **Deploy guard.** If the deployed gateway commit changed during a pair, the pair was invalid and ran again.
- **Cost.** List price applied to the token counts Anthropic returned for each call. The runs used a claude.ai login, which has no per-token invoice. Direct: Claude Code's own total. Context Mode: the gateway's per-call record, hidden calls included. See [Cost basis](https://context-mode.com/docs/how-saving-works#cost-basis).
- **Analysis.** Every valid pair, no outlier dropped. Change = mean(Context Mode minus direct) divided by mean(direct), with the 95% t interval. "Less" when the whole interval is below zero, "more" when above, else "same". Wins count the pairs where Context Mode cost less.

## Claude Opus 5.5

| Workload | Pairs | Change (95% range) | Verdict | Wins | Success, direct and Context Mode | Gateway commit | Evidence |
| --- | --- | --- | --- | --- | --- | --- | --- |
| Three parallel subagents, then one more | 20 | −15.2% (−18.7 to −11.7) | less | 20 of 20 | 20/20, 20/20 | 3cba9400 | [JSON](https://context-mode.com/benchmarks/data/v4-opus55-W1-W2-W3-W4.json) · [Pre-registration](https://context-mode.com/benchmarks/prereg/v4-opus55-prereg) |
| Large documents processed with a script | 20 | −9.9% (−15.4 to −4.4) | less | 15 of 20 | 20/20, 20/20 | 3cba9400 | [JSON](https://context-mode.com/benchmarks/data/v4-opus55-W1-W2-W3-W4.json) · [Pre-registration](https://context-mode.com/benchmarks/prereg/v4-opus55-prereg) |
| `npm test` fix loop, short | 16 | −21.8% (−27.7 to −15.8) | less | 15 of 16 | 16/16, 16/16 | 408333a4 | [JSON](https://context-mode.com/benchmarks/data/v3-opus55-short.json) · [Pre-registration](https://context-mode.com/benchmarks/prereg/v3-opus55-short-prereg) |
| `npm test` fix loop, long | 12 | −21.3% (−28.3 to −14.2) | less | 11 of 12 | 12/12, 12/12 | 3bedf1ba | [JSON](https://context-mode.com/benchmarks/data/final-opus55-tests-W3-minitest-W4.json) · [Pre-registration](https://context-mode.com/benchmarks/prereg/final-opus55-tests-W3-minitest-W4-prereg) |
| Test fix loop (minitest) | 8 | −61.0% (−69.9 to −52.2) | less | 8 of 8 | 8/8, 8/8 | 3bedf1ba | [JSON](https://context-mode.com/benchmarks/data/final-opus55-tests-W3-minitest-W4.json) · [Pre-registration](https://context-mode.com/benchmarks/prereg/final-opus55-tests-W3-minitest-W4-prereg) |
| One large log | 16 | −22.4% (−30.2 to −14.7) | less | 15 of 16 | 16/16, 16/16 | 408333a4 | [JSON](https://context-mode.com/benchmarks/data/v3-opus55-short.json) · [Pre-registration](https://context-mode.com/benchmarks/prereg/v3-opus55-short-prereg) |
| One large JSON file | 12 | −21.4% (−29.9 to −12.9) | less | 12 of 12 | 12/12, 12/12 | 408333a4 | [JSON](https://context-mode.com/benchmarks/data/v3-opus55-short.json) · [Pre-registration](https://context-mode.com/benchmarks/prereg/v3-opus55-short-prereg) |
| Repo history question ( `git log` , 320 commits) | 16 | −25.2% (−33.9 to −16.5) | less | 15 of 16 | 16/16, 16/16 | 408333a4 | [JSON](https://context-mode.com/benchmarks/data/v3-opus55-short.json) · [Pre-registration](https://context-mode.com/benchmarks/prereg/v3-opus55-short-prereg) |
| Code search ( `grep` over 150 files) | 16 | −28.7% (−34.1 to −23.4) | less | 16 of 16 | 16/16, 16/16 | 408333a4 | [JSON](https://context-mode.com/benchmarks/data/v3-opus55-short.json) · [Pre-registration](https://context-mode.com/benchmarks/prereg/v3-opus55-short-prereg) |
| One-line edit | 12 | −27.8% (−42.2 to −13.5) | less | 11 of 12 | 12/12, 12/12 | 408333a4 | [JSON](https://context-mode.com/benchmarks/data/v3-opus55-short.json) · [Pre-registration](https://context-mode.com/benchmarks/prereg/v3-opus55-short-prereg) |

Rows where cost was lower on Opus 5.5.

## Claude Haiku 4.5

| Workload | Pairs | Change (95% range) | Verdict | Wins | Success, direct and Context Mode | Gateway commit | Evidence |
| --- | --- | --- | --- | --- | --- | --- | --- |
| `npm test` fix loop | 12 | −52.8% (−56.2 to −49.4) | less | 12 of 12 | 12/12, 12/12 | 3bedf1ba | [JSON](https://context-mode.com/benchmarks/data/final-haiku45.json) · [Pre-registration](https://context-mode.com/benchmarks/prereg/final-haiku45-prereg) |
| Test fix loop (minitest) | 8 | −45.5% (−56.8 to −34.2) | less | 8 of 8 | 8/8, 8/8 | 3bedf1ba | [JSON](https://context-mode.com/benchmarks/data/final-haiku45.json) · [Pre-registration](https://context-mode.com/benchmarks/prereg/final-haiku45-prereg) |
| Large documents processed with a script | 12 | −64.4% (−76.5 to −52.3) | less | 12 of 12 | 11/12, 11/12 | 3bedf1ba | [JSON](https://context-mode.com/benchmarks/data/final-haiku45.json) · [Pre-registration](https://context-mode.com/benchmarks/prereg/final-haiku45-prereg) |
| Repo history question ( `git log` ) | 12 | −19.2% (−22.5 to −15.8) | less | 12 of 12 | 12/12, 12/12 | 3bedf1ba | [JSON](https://context-mode.com/benchmarks/data/final-haiku45.json) · [Pre-registration](https://context-mode.com/benchmarks/prereg/final-haiku45-prereg) |
| Code search ( `grep` ) | 20 | −28.9% (−53.8 to −3.9) | less | 18 of 20 | 20/20, 20/20 | 3bedf1ba | [JSON](https://context-mode.com/benchmarks/data/final-haiku45.json) · [Pre-registration](https://context-mode.com/benchmarks/prereg/final-haiku45-prereg) |
| One large JSON file | 20 | −16.1% (−22.8 to −9.3) | less | 18 of 20 | 20/20, 20/20 | 3bedf1ba | [JSON](https://context-mode.com/benchmarks/data/final-haiku45.json) · [Pre-registration](https://context-mode.com/benchmarks/prereg/final-haiku45-prereg) |
| One-line edit | 8 | −20.0% (−24.5 to −15.4) | less | 8 of 8 | 8/8, 8/8 | 3bedf1ba | [JSON](https://context-mode.com/benchmarks/data/final-haiku45.json) · [Pre-registration](https://context-mode.com/benchmarks/prereg/final-haiku45-prereg) |

Rows where cost was lower on Haiku 4.5. Not run on Haiku 4.5: multi-agent.

## Evidence files

Each JSON lists every counted pair with its cost per arm at list price, success and turns, and the analysis recomputed from those pairs. Prompts, answers, session and tenant ids, hostnames and local paths are removed. The [evidence index](https://context-mode.com/benchmarks/data/index.json) states the method and every removed field. We ran no Codex cost test and quote no Codex figure.

[Back to benchmarks](https://context-mode.com/docs/benchmarks)
