What is Context Mode Gateway?

Context Mode Gateway is a gateway between Claude Code or Codex and the model API. It folds tool output the first time the model sees it, checks tool calls against your rules, and shows what every request cost.

Key facts

What it is
A hosted gateway on the request path of your coding agent, set with one base URL per agent. The first of three products: Context Mode Gateway, Context Mode Engine and Context Mode Insight.
Who it is for
Developers and teams who run Claude Code or Codex and want lower cost, one tool-call policy and one memory for both.
Clients
Claude Code and Codex. Anthropic and OpenAI models, paid by your own Claude or OpenAI account.
Measured result
Claude Code on Claude Opus 5.5 cost 65.5% less on a long agentic coding session, which went from $7.20 to $2.49, 4 of 4 pairs cheaper, run on 2026-10-03. Over 168 shorter paired runs on 2026-09-29, 155 cost less and 168 of 168 tasks passed in both arms. Method and evidence.
Limits
Runs on Cloudflare, with no on-premises build. No SSO or team roles yet. Codex cost is not measured yet. Its own tools, memory and skills add tokens, and the benchmarks count them.
01 your agentClaude Code or Codex, with its base URL set to gateway.context-mode.com.
02 gatewayFolds tool output, trims tool schemas, adds memory and skills, checks Cage rules, records tokens and cost.
03 model APIAnthropic or OpenAI, called with your own credential. Your provider bills you as before.
04 consoleconsole.context-mode.com shows each request's tokens, cost and saving.

What it saves

Our latest pre-registered paired runs against Claude Code going direct, on Claude Opus 5.5: everyday tasks on 2026-09-29, long sessions on 2026-10-03 and 2026-10-04. Task success was equal in both arms on every row.

CategoryTaskPairsGateway vs direct
Agentic CodingCode search (grep over 150 files)16Opus 5.5: 29% lower cost (23 to 34%)
Agentic CodingOne-line edit12Opus 5.5: 28% lower cost (14 to 42%)
Agentic CodingRepo history question (git log, 320 commits)16Opus 5.5: 25% lower cost (16 to 34%)
Large Tool-Output HandlingOne large log16Opus 5.5: 22% lower cost (15 to 30%)
Large Tool-Output HandlingOne large JSON file12Opus 5.5: 21% lower cost (13 to 30%)
Test-Driven Debuggingnpm test fix loop, short16Opus 5.5: 22% lower cost (16 to 28%)
Test-Driven Debuggingnpm test fix loop, long12Opus 5.5: 21% lower cost (14 to 28%)
Multi-Agent OrchestrationThree parallel subagents, then one more20Opus 5.5: 15% lower cost (12 to 19%)
Long Agentic Coding SessionLong agentic coding session, Opus 5.5, 1M window4Opus 5.5: 65.5% lower cost (64.5 to 67.7%, lowest to highest pair), $7.20 to $2.49
Long Agentic Coding SessionLong agentic coding session, Opus 5.5, 200K window, Recall on3Opus 5.5: 37.6% lower cost (30.7 to 44.3%, lowest to highest pair), $4.86 to $3.04; 9 of 10 exact answers about early details against 4 of 10
Long-Context Data AnalysisLarge documents processed with a script20Opus 5.5: 10% lower cost (4 to 15%)

The gateway saves where tool output is large or repeats. Use only Context Saving, with Cage, Memory, Skills and Recall off, and the everyday-task rows show your saving: those runs had Cage and Recall off, and the tokens Memory added were counted against the gateway. Every result for both models, with its method and evidence.

The feature map

Context SavingTest, build, install and log output folded the first time the model sees it. The full output is archived. Docs
Thinking in CodeThe agent writes a short script. The gateway runs it in a sandbox, or on your machine for local files, and only what it prints reaches the model. Docs
CageOff until you turn it on. Balanced and Locked down stop commands that wipe a disk or delete your home folder, and your rules decide which sites, commands and scripts the agent may reach. Docs
RecallA long Claude Code or Codex session that never compacts: the model gets recent work and an index, and the agent fetches any older message or tool output from an archive, word for word except masked secrets and e-mail addresses. Docs
MemoryFacts from your turns, kept with two clocks and added on the first request of a conversation, and after the client summarises the conversation. It adds tokens to that request. Docs
SkillsImport skills from GitHub. The gateway adds the matching skill to the prompt. Docs
SearchPrompts, commands and archived output from every session. Archived output reads back byte for byte, except secrets and e-mail addresses, which are masked before storage. Docs
Context limitHow much history a thread keeps: 200K, 500K or 1M, set once in Settings. Docs
Reply styleOne writing style for every reply, in Claude Code and Codex. Docs
Session resumeOpen a Claude Code or Codex session on another machine, in either client. Docs
ConsoleTokens, cost and saving per request, by session, project and agent.

What changes

What does not change

Plans

Start free: 2,000 requests, every feature, no card. Every plan has every feature, in Claude Code and Codex. Pro gives 50,000 requests a month. Team gives each seat everything in Pro, with one pool of requests, one Cage policy for the org and one bill. When your requests run out, nothing breaks: your requests go straight to the model. Plans and limits.

Where it fits

How does Context Mode compare with other tools? See the landscape: gateways, token optimizers, agent firewalls, memory layers and more.

Context Mode Engine and Insight

Context Mode has three products. The Gateway, on this page, is the main one. Context Mode Engine is the free context-mode plugin, where it started: it keeps tool output out of one agent's context window, on your machine, in 18 AI coding tools, with no account. Context Mode Insight reads the Engine's events and shows how a team works with coding agents.

FAQ

Is Context Mode a plugin or a gateway?

Both exist. Context Mode Gateway, on this page, is the hosted gateway. Context Mode Engine is the free context-mode plugin that runs on your machine.

Do you hold my Claude or OpenAI credential?

No. The gateway forwards it with each request, and we do not resell tokens.

Does the gateway see my prompts?

Yes. It sits on the request path and reads every request. What we store.

Which model providers does it support?

Anthropic and OpenAI only.

How do I leave?

Remove the lines the CLI added. Your agent then talks to the model API directly. Steps.