Context Mode launch benchmark: the full table

The workloads measured on 2026-09-29 where cost was lower, each with a published evidence file. This is not every workload we ran: the pre-registrations marked "edited" show where workloads were taken off. Each row is a paired A/B test: the same task, run with Claude Code direct to Anthropic and through Context Mode, back to back.

Method

Claude Opus 5.5

WorkloadPairsChange (95% range)VerdictWinsSuccess, direct and Context ModeGateway commitEvidence
Three parallel subagents, then one more20−15.2% (−18.7 to −11.7)less20 of 2020/20, 20/203cba9400JSON · Pre-registration
Large documents processed with a script20−9.9% (−15.4 to −4.4)less15 of 2020/20, 20/203cba9400JSON · Pre-registration
npm test fix loop, short16−21.8% (−27.7 to −15.8)less15 of 1616/16, 16/16408333a4JSON · Pre-registration
npm test fix loop, long12−21.3% (−28.3 to −14.2)less11 of 1212/12, 12/123bedf1baJSON · Pre-registration
Test fix loop (minitest)8−61.0% (−69.9 to −52.2)less8 of 88/8, 8/83bedf1baJSON · Pre-registration
One large log16−22.4% (−30.2 to −14.7)less15 of 1616/16, 16/16408333a4JSON · Pre-registration
One large JSON file12−21.4% (−29.9 to −12.9)less12 of 1212/12, 12/12408333a4JSON · Pre-registration
Repo history question (git log, 320 commits)16−25.2% (−33.9 to −16.5)less15 of 1616/16, 16/16408333a4JSON · Pre-registration
Code search (grep over 150 files)16−28.7% (−34.1 to −23.4)less16 of 1616/16, 16/16408333a4JSON · Pre-registration
One-line edit12−27.8% (−42.2 to −13.5)less11 of 1212/12, 12/12408333a4JSON · Pre-registration

Rows where cost was lower on Opus 5.5.

Claude Haiku 4.5

WorkloadPairsChange (95% range)VerdictWinsSuccess, direct and Context ModeGateway commitEvidence
npm test fix loop12−52.8% (−56.2 to −49.4)less12 of 1212/12, 12/123bedf1baJSON · Pre-registration
Test fix loop (minitest)8−45.5% (−56.8 to −34.2)less8 of 88/8, 8/83bedf1baJSON · Pre-registration
Large documents processed with a script12−64.4% (−76.5 to −52.3)less12 of 1211/12, 11/123bedf1baJSON · Pre-registration
Repo history question (git log)12−19.2% (−22.5 to −15.8)less12 of 1212/12, 12/123bedf1baJSON · Pre-registration
Code search (grep)20−28.9% (−53.8 to −3.9)less18 of 2020/20, 20/203bedf1baJSON · Pre-registration
One large JSON file20−16.1% (−22.8 to −9.3)less18 of 2020/20, 20/203bedf1baJSON · Pre-registration
One-line edit8−20.0% (−24.5 to −15.4)less8 of 88/8, 8/83bedf1baJSON · Pre-registration

Rows where cost was lower on Haiku 4.5. Not run on Haiku 4.5: multi-agent.

Evidence files

Each JSON lists every counted pair with its cost per arm at list price, success and turns, and the analysis recomputed from those pairs. Prompts, answers, session and tenant ids, hostnames and local paths are removed. The evidence index states the method and every removed field. We ran no Codex cost test and quote no Codex figure.