# What is Context Mode Gateway?

Context Mode Gateway is a gateway between Claude Code or Codex and the model API. It folds tool output the first time the model sees it, checks tool calls against your rules, and shows what every request cost.

## Key facts

What it is

A hosted gateway on the request path of your coding agent, set with one base URL per agent. The first of three products: Context Mode Gateway, Context Mode Engine and Context Mode Insight.

Who it is for

Developers and teams who run Claude Code or Codex and want lower cost, one tool-call policy and one memory for both.

Clients

Claude Code and Codex. Anthropic and OpenAI models, paid by your own Claude or OpenAI account.

Measured result

Claude Code on Claude Opus 5.5 cost [65.5% less](https://context-mode.com/benchmarks/long-session/ab.json) on a long agentic coding session, which went from [$7.20](https://context-mode.com/benchmarks/long-session/ab.json) to [$2.49](https://context-mode.com/benchmarks/long-session/ab.json), 4 of 4 pairs cheaper, run on 2026-10-03. Over 168 shorter paired runs on 2026-09-29, 155 cost less and 168 of 168 tasks passed in both arms. [Method and evidence](https://context-mode.com/docs/benchmarks#long-session).

Limits

Runs on Cloudflare, with no on-premises build. No SSO or team roles yet. Codex cost is not measured yet. Its own tools, memory and skills add tokens, and the benchmarks count them.

01 your agent Claude Code or Codex, with its base URL set to **gateway.context-mode.com**.

02 gateway Folds tool output, trims tool schemas, adds memory and skills, checks Cage rules, records tokens and cost.

03 model API Anthropic or OpenAI, called with **your own credential**. Your provider bills you as before.

04 console **console.context-mode.com** shows each request's tokens, cost and saving.

## What it saves

Our latest pre-registered paired runs against Claude Code going direct, on Claude Opus 5.5: everyday tasks on 2026-09-29, long sessions on 2026-10-03 and 2026-10-04. Task success was equal in both arms on every row.

| Category | Task | Pairs | Gateway vs direct |
| --- | --- | --- | --- |
| [Agentic Coding](https://context-mode.com/docs/benchmarks#agentic-coding) | Code search ( `grep` over 150 files) | 16 | Opus 5.5: 29% lower cost (23 to 34%) |
| [Agentic Coding](https://context-mode.com/docs/benchmarks#agentic-coding) | One-line edit | 12 | Opus 5.5: 28% lower cost (14 to 42%) |
| [Agentic Coding](https://context-mode.com/docs/benchmarks#agentic-coding) | Repo history question ( `git log` , 320 commits) | 16 | Opus 5.5: 25% lower cost (16 to 34%) |
| [Large Tool-Output Handling](https://context-mode.com/docs/benchmarks#tool-output) | One large log | 16 | Opus 5.5: 22% lower cost (15 to 30%) |
| [Large Tool-Output Handling](https://context-mode.com/docs/benchmarks#tool-output) | One large JSON file | 12 | Opus 5.5: 21% lower cost (13 to 30%) |
| [Test-Driven Debugging](https://context-mode.com/docs/benchmarks#test-debugging) | `npm test` fix loop, short | 16 | Opus 5.5: 22% lower cost (16 to 28%) |
| [Test-Driven Debugging](https://context-mode.com/docs/benchmarks#test-debugging) | `npm test` fix loop, long | 12 | Opus 5.5: 21% lower cost (14 to 28%) |
| [Multi-Agent Orchestration](https://context-mode.com/docs/benchmarks#multi-agent) | Three parallel subagents, then one more | 20 | Opus 5.5: 15% lower cost (12 to 19%) |
| [Long Agentic Coding Session](https://context-mode.com/docs/benchmarks#long-session) | Long agentic coding session, Opus 5.5, 1M window | 4 | Opus 5.5: [65.5% lower cost](https://context-mode.com/benchmarks/long-session/ab.json) (64.5 to 67.7%, lowest to highest pair), [$7.20](https://context-mode.com/benchmarks/long-session/ab.json) to [$2.49](https://context-mode.com/benchmarks/long-session/ab.json) |
| [Long Agentic Coding Session](https://context-mode.com/docs/benchmarks#long-session) | Long agentic coding session, Opus 5.5, 200K window, Recall on | 3 | Opus 5.5: [37.6% lower cost](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) ( [30.7 to 44.3%](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) , lowest to highest pair), [$4.86](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) to [$3.04](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) ; [9 of 10](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) exact answers about early details against [4 of 10](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) |
| [Long-Context Data Analysis](https://context-mode.com/docs/benchmarks#data-analysis) | Large documents processed with a script | 20 | Opus 5.5: 10% lower cost (4 to 15%) |

The gateway saves where tool output is large or repeats. Use only Context Saving, with Cage, Memory, Skills and Recall off, and the everyday-task rows show your saving: those runs had Cage and Recall off, and the tokens Memory added were counted against the gateway. [Every result for both models, with its method and evidence](https://context-mode.com/docs/benchmarks).

## The feature map

[Context Saving](https://context-mode.com/context-saving) Test, build, install and log output folded the first time the model sees it. The full output is archived. [Docs](https://context-mode.com/docs/how-saving-works)

[Thinking in Code](https://context-mode.com/thinking-in-code) The agent writes a short script. The gateway runs it in a sandbox, or on your machine for local files, and only what it prints reaches the model. [Docs](https://context-mode.com/docs/how-saving-works#code-mode)

[Cage](https://context-mode.com/cage) Off until you turn it on. Balanced and Locked down stop commands that wipe a disk or delete your home folder, and your rules decide which sites, commands and scripts the agent may reach. [Docs](https://context-mode.com/docs/cage)

[Recall](https://context-mode.com/recall) A long Claude Code or Codex session that never compacts: the model gets recent work and an index, and the agent fetches any older message or tool output from an archive, word for word except masked secrets and e-mail addresses. [Docs](https://context-mode.com/docs/recall)

[Memory](https://context-mode.com/memory) Facts from your turns, kept with two clocks and added on the first request of a conversation, and after the client summarises the conversation. It adds tokens to that request. [Docs](https://context-mode.com/docs/memory)

[Skills](https://context-mode.com/skills) Import skills from GitHub. The gateway adds the matching skill to the prompt. [Docs](https://context-mode.com/docs/skills)

[Search](https://context-mode.com/search) Prompts, commands and archived output from every session. Archived output reads back byte for byte, except secrets and e-mail addresses, which are masked before storage. [Docs](https://context-mode.com/docs/search)

[Context limit](https://context-mode.com/docs/context-limit) How much history a thread keeps: 200K, 500K or 1M, set once in Settings. [Docs](https://context-mode.com/docs/context-limit)

[Reply style](https://context-mode.com/docs/reply) One writing style for every reply, in Claude Code and Codex. [Docs](https://context-mode.com/docs/reply)

[Session resume](https://context-mode.com/session-resume) Open a Claude Code or Codex session on another machine, in either client. [Docs](https://context-mode.com/docs/session-resume)

[Console](https://console.context-mode.com) Tokens, cost and saving per request, by session, project and agent.

## What changes

- Tool output is folded once and sent the same way on every later turn, so the prompt cache keeps reading it.
- Tool schemas are trimmed. The gateway adds its own tools, memory and skills, and those cost tokens. The benchmarks count both sides.
- Cage, once you turn it on, checks each tool call. A blocked command never runs: your terminal runs only an echo that carries the refusal.
- Every request is recorded: tokens, cost and saving.

## What does not change

- Your agent, your model and your subscription. The gateway forwards your credential, and we do not resell tokens.
- No plugin or background service is installed. The CLI edits one config file per agent and keeps a backup.
- Nothing is cut silently. A folded result ends with a line that says what was dropped and how to read the rest.
- To leave, [remove the lines the CLI added](https://context-mode.com/docs/faq#leave). Your agent then talks to the model API directly.

## Plans

Start free: 2,000 requests, every feature, no card. Every plan has every feature, in Claude Code and Codex. Pro gives 50,000 requests a month. Team gives each seat everything in Pro, with one pool of requests, one Cage policy for the org and one bill. When your requests run out, nothing breaks: your requests go straight to the model. [Plans and limits](https://context-mode.com/docs/plans-and-limits).

## Where it fits

How does Context Mode compare with other tools? See the [landscape](https://context-mode.com/docs/landscape): gateways, token optimizers, agent firewalls, memory layers and more.

## Context Mode Engine and Insight

Context Mode has three products. The Gateway, on this page, is the main one. [Context Mode Engine](https://context-mode.com/engine) is the free context-mode plugin, where it started: it keeps tool output out of one agent's context window, on your machine, in 18 AI coding tools, with no account. [Context Mode Insight](https://context-mode.com/insight) reads the Engine's events and shows how a team works with coding agents.

## FAQ

### Is Context Mode a plugin or a gateway?

Both exist. Context Mode Gateway, on this page, is the hosted gateway. [Context Mode Engine](https://context-mode.com/engine) is the free context-mode plugin that runs on your machine.

### Do you hold my Claude or OpenAI credential?

No. The gateway forwards it with each request, and we do not resell tokens.

### Does the gateway see my prompts?

Yes. It sits on the request path and reads every request. [What we store](https://context-mode.com/docs/faq#store).

### Which model providers does it support?

Anthropic and OpenAI only.

### How do I leave?

Remove the lines the CLI added. Your agent then talks to the model API directly. [Steps](https://context-mode.com/docs/faq#leave).

[Connect your agent](https://context-mode.com/docs/quick-start) [See the benchmarks](https://context-mode.com/docs/benchmarks) [Open the console](https://console.context-mode.com)
