What is a coding agent gateway?
A coding agent gateway is a service between your coding agent, such as Claude Code or Codex, and the model API that shortens tool output and checks tool calls, not only routes and logs them. You use one by pointing your agent's base URL at it, and Context Mode Gateway is one for Claude Code and Codex.
What to do:
- Check that you need one: see when you need one.
- Connect Claude Code or Codex to Context Mode Gateway with one command:
npx @context-mode/cli. See the quick start. - Watch your own tokens, cost and saving per request in the console.
Key facts
- What it is
- A proxy set as the agent's base URL. Context Mode Gateway is one: it folds tool output the first time the model sees it, checks tool calls against your rules, and shows what every request cost.
- Who it is for
- CTOs, platform engineers and developers who run Claude Code or Codex and want to change what the agent sends and runs, not only where it is routed.
- Clients
- Context Mode Gateway: Claude Code and Codex, with Anthropic and OpenAI models paid by your own account.
- Measured result
- Claude Code on Claude Opus 5.5 cost 65.5% less on a long agentic coding session, which went from $7.20 to $2.49, 4 of 4 pairs cheaper, run on 2026-10-03. Over 168 shorter paired runs on 2026-09-29, 155 cost less and 168 of 168 tasks passed in both arms. Method and evidence.
- Limits
- Every prompt passes the gateway. Context Mode Gateway is hosted only, routes to Anthropic and OpenAI only, and has no SSO or team roles yet.
How an agent reaches a gateway
Claude Code reads its base URL from ANTHROPIC_BASE_URL. Anthropic documents routing Claude Code through a gateway your organization runs, and says a gateway gives one place for credentials, usage tracking, cost controls, audit logging and provider switching (Other LLM gateways, read 2026-10-01). Codex takes a custom model provider in ~/.codex/config.toml. The Context Mode CLI writes both, with a backup of each file.
What a coding agent gateway adds
A coding agent spends most of its tokens on tool output and tool schemas, and its risk is in the tool calls it asks the client to run. A gateway that only routes and logs leaves all three as they are. The table puts Context Mode Gateway next to the general gateway Anthropic's page describes.
| Question | Context Mode Gateway | General LLM gateway, as Anthropic describes it |
|---|---|---|
| Where it sits | Between Claude Code or Codex and the model API | Between Claude Code and your provider |
| Credentials | Your own Claude or OpenAI credential, forwarded with each request | The provider key stays on the gateway; developers hold gateway credentials |
| Tool output | Folded the first time it passes, full output archived | Not on Anthropic's list |
| Tool calls | Cage checks each call before the client runs it | Not on Anthropic's list |
| Usage and cost | Tokens, cost and saving per request, by session, project and agent | Usage by developer or team, budgets and rate limits |
| Model providers | Anthropic and OpenAI | Provider switching in the gateway's configuration |
| Who runs it | Hosted by us | Your organization |
For named gateways, token optimizers and agent firewalls, with sources, see the landscape.
When you need one
- Cost. Your agents run tests, read logs and large files, or start subagents, and you want that output paid for once. See How to reduce Claude Code costs.
- One policy. You run Claude Code and Codex and want one set of tool-call rules for both. See Codex in the quick start.
- One memory and history. You want facts, skills, search and session resume to follow you across machines and clients.
You may not need one if you run a single agent for short questions with little tool output. There is little to save, and the client's own permission rules and sandbox may be enough.
FAQ
How is it different from a general LLM gateway?
A general LLM gateway routes many models through one API and logs the calls. Context Mode Gateway routes to Anthropic and OpenAI only, and works on what the agent reads and runs.
Is it an MCP gateway?
No. Context Mode Gateway is a gateway for the model API, not for MCP servers. Cage rules also cover calls to tool servers, so one policy covers shell commands, sites, files and MCP tools.
Does the gateway see my prompts?
Yes. It sits on the request path and reads every request. See what we store.
Do I need one if I use one agent?
It also gives that one agent search, session resume and per-request cost. Whether it saves money depends on how much tool output your work makes.