What is a coding agent gateway?

A coding agent gateway is a service between your coding agent, such as Claude Code or Codex, and the model API that shortens tool output and checks tool calls, not only routes and logs them. You use one by pointing your agent's base URL at it, and Context Mode Gateway is one for Claude Code and Codex.

What to do:

  1. Check that you need one: see when you need one.
  2. Connect Claude Code or Codex to Context Mode Gateway with one command: npx @context-mode/cli. See the quick start.
  3. Watch your own tokens, cost and saving per request in the console.

Key facts

What it is
A proxy set as the agent's base URL. Context Mode Gateway is one: it folds tool output the first time the model sees it, checks tool calls against your rules, and shows what every request cost.
Who it is for
CTOs, platform engineers and developers who run Claude Code or Codex and want to change what the agent sends and runs, not only where it is routed.
Clients
Context Mode Gateway: Claude Code and Codex, with Anthropic and OpenAI models paid by your own account.
Measured result
Claude Code on Claude Opus 5.5 cost 65.5% less on a long agentic coding session, which went from $7.20 to $2.49, 4 of 4 pairs cheaper, run on 2026-10-03. Over 168 shorter paired runs on 2026-09-29, 155 cost less and 168 of 168 tasks passed in both arms. Method and evidence.
Limits
Every prompt passes the gateway. Context Mode Gateway is hosted only, routes to Anthropic and OpenAI only, and has no SSO or team roles yet.

How an agent reaches a gateway

Claude Code reads its base URL from ANTHROPIC_BASE_URL. Anthropic documents routing Claude Code through a gateway your organization runs, and says a gateway gives one place for credentials, usage tracking, cost controls, audit logging and provider switching (Other LLM gateways, read 2026-10-01). Codex takes a custom model provider in ~/.codex/config.toml. The Context Mode CLI writes both, with a backup of each file.

What a coding agent gateway adds

A coding agent spends most of its tokens on tool output and tool schemas, and its risk is in the tool calls it asks the client to run. A gateway that only routes and logs leaves all three as they are. The table puts Context Mode Gateway next to the general gateway Anthropic's page describes.

QuestionContext Mode GatewayGeneral LLM gateway, as Anthropic describes it
Where it sitsBetween Claude Code or Codex and the model APIBetween Claude Code and your provider
CredentialsYour own Claude or OpenAI credential, forwarded with each requestThe provider key stays on the gateway; developers hold gateway credentials
Tool outputFolded the first time it passes, full output archivedNot on Anthropic's list
Tool callsCage checks each call before the client runs itNot on Anthropic's list
Usage and costTokens, cost and saving per request, by session, project and agentUsage by developer or team, budgets and rate limits
Model providersAnthropic and OpenAIProvider switching in the gateway's configuration
Who runs itHosted by usYour organization

For named gateways, token optimizers and agent firewalls, with sources, see the landscape.

When you need one

You may not need one if you run a single agent for short questions with little tool output. There is little to save, and the client's own permission rules and sandbox may be enough.

FAQ

How is it different from a general LLM gateway?

A general LLM gateway routes many models through one API and logs the calls. Context Mode Gateway routes to Anthropic and OpenAI only, and works on what the agent reads and runs.

Is it an MCP gateway?

No. Context Mode Gateway is a gateway for the model API, not for MCP servers. Cage rules also cover calls to tool servers, so one policy covers shell commands, sites, files and MCP tools.

Does the gateway see my prompts?

Yes. It sits on the request path and reads every request. See what we store.

Do I need one if I use one agent?

It also gives that one agent search, session resume and per-request cost. Whether it saves money depends on how much tool output your work makes.