What is Context Mode Gateway?
Context Mode Gateway is a gateway between Claude Code or Codex and the model API. It folds tool output the first time the model sees it, checks tool calls against your rules, and shows what every request cost.
Key facts
- What it is
- A hosted gateway on the request path of your coding agent, set with one base URL per agent. The first of three products: Context Mode Gateway, Context Mode Engine and Context Mode Insight.
- Who it is for
- Developers and teams who run Claude Code or Codex and want lower cost, one tool-call policy and one memory for both.
- Clients
- Claude Code and Codex. Anthropic and OpenAI models, paid by your own Claude or OpenAI account.
- Measured result
- Claude Code on Claude Opus 5.5 cost 65.5% less on a long agentic coding session, which went from $7.20 to $2.49, 4 of 4 pairs cheaper, run on 2026-10-03. Over 168 shorter paired runs on 2026-09-29, 155 cost less and 168 of 168 tasks passed in both arms. Method and evidence.
- Limits
- Runs on Cloudflare, with no on-premises build. No SSO or team roles yet. Codex cost is not measured yet. Its own tools, memory and skills add tokens, and the benchmarks count them.
What it saves
Our latest pre-registered paired runs against Claude Code going direct, on Claude Opus 5.5: everyday tasks on 2026-09-29, long sessions on 2026-10-03 and 2026-10-04. Task success was equal in both arms on every row.
| Category | Task | Pairs | Gateway vs direct |
|---|---|---|---|
| Agentic Coding | Code search (grep over 150 files) | 16 | Opus 5.5: 29% lower cost (23 to 34%) |
| Agentic Coding | One-line edit | 12 | Opus 5.5: 28% lower cost (14 to 42%) |
| Agentic Coding | Repo history question (git log, 320 commits) | 16 | Opus 5.5: 25% lower cost (16 to 34%) |
| Large Tool-Output Handling | One large log | 16 | Opus 5.5: 22% lower cost (15 to 30%) |
| Large Tool-Output Handling | One large JSON file | 12 | Opus 5.5: 21% lower cost (13 to 30%) |
| Test-Driven Debugging | npm test fix loop, short | 16 | Opus 5.5: 22% lower cost (16 to 28%) |
| Test-Driven Debugging | npm test fix loop, long | 12 | Opus 5.5: 21% lower cost (14 to 28%) |
| Multi-Agent Orchestration | Three parallel subagents, then one more | 20 | Opus 5.5: 15% lower cost (12 to 19%) |
| Long Agentic Coding Session | Long agentic coding session, Opus 5.5, 1M window | 4 | Opus 5.5: 65.5% lower cost (64.5 to 67.7%, lowest to highest pair), $7.20 to $2.49 |
| Long Agentic Coding Session | Long agentic coding session, Opus 5.5, 200K window, Recall on | 3 | Opus 5.5: 37.6% lower cost (30.7 to 44.3%, lowest to highest pair), $4.86 to $3.04; 9 of 10 exact answers about early details against 4 of 10 |
| Long-Context Data Analysis | Large documents processed with a script | 20 | Opus 5.5: 10% lower cost (4 to 15%) |
The gateway saves where tool output is large or repeats. Use only Context Saving, with Cage, Memory, Skills and Recall off, and the everyday-task rows show your saving: those runs had Cage and Recall off, and the tokens Memory added were counted against the gateway. Every result for both models, with its method and evidence.
The feature map
What changes
- Tool output is folded once and sent the same way on every later turn, so the prompt cache keeps reading it.
- Tool schemas are trimmed. The gateway adds its own tools, memory and skills, and those cost tokens. The benchmarks count both sides.
- Cage, once you turn it on, checks each tool call. A blocked command never runs: your terminal runs only an echo that carries the refusal.
- Every request is recorded: tokens, cost and saving.
What does not change
- Your agent, your model and your subscription. The gateway forwards your credential, and we do not resell tokens.
- No plugin or background service is installed. The CLI edits one config file per agent and keeps a backup.
- Nothing is cut silently. A folded result ends with a line that says what was dropped and how to read the rest.
- To leave, remove the lines the CLI added. Your agent then talks to the model API directly.
Plans
Start free: 2,000 requests, every feature, no card. Every plan has every feature, in Claude Code and Codex. Pro gives 50,000 requests a month. Team gives each seat everything in Pro, with one pool of requests, one Cage policy for the org and one bill. When your requests run out, nothing breaks: your requests go straight to the model. Plans and limits.
Where it fits
How does Context Mode compare with other tools? See the landscape: gateways, token optimizers, agent firewalls, memory layers and more.
Context Mode Engine and Insight
Context Mode has three products. The Gateway, on this page, is the main one. Context Mode Engine is the free context-mode plugin, where it started: it keeps tool output out of one agent's context window, on your machine, in 18 AI coding tools, with no account. Context Mode Insight reads the Engine's events and shows how a team works with coding agents.
FAQ
Is Context Mode a plugin or a gateway?
Both exist. Context Mode Gateway, on this page, is the hosted gateway. Context Mode Engine is the free context-mode plugin that runs on your machine.
Do you hold my Claude or OpenAI credential?
No. The gateway forwards it with each request, and we do not resell tokens.
Does the gateway see my prompts?
Yes. It sits on the request path and reads every request. What we store.
Which model providers does it support?
Anthropic and OpenAI only.
How do I leave?
Remove the lines the CLI added. Your agent then talks to the model API directly. Steps.