FAQ about Context Mode Gateway

Context Mode Gateway is a gateway between Claude Code or Codex and the model API, and your own Claude or OpenAI account pays the model. Below are straight answers on cost, data and control; where we do not have something yet, we say so.

Key facts

What it is
A hosted gateway between Claude Code or Codex and the model API. It folds tool output, checks tool calls against your rules and records what every request cost.
Who it is for
Developers and teams who run Claude Code or Codex.
Clients
Claude Code and Codex. Anthropic and OpenAI models, paid by your own Claude or OpenAI account.
Measured result
Claude Code on Claude Opus 5.5 cost 65.5% less on a long agentic coding session, which went from $7.20 to $2.49, 4 of 4 pairs cheaper, run on 2026-10-03. Over 168 shorter paired runs on 2026-09-29, 155 cost less and 168 of 168 tasks passed in both arms. Method and evidence.
Limits
Runs on Cloudflare, with no on-premises build. No SSO or team roles yet. Codex cost is not measured yet.

Cost

How do I get Context Mode Gateway?

Run npx @context-mode/cli, or sign in to the console. Start free: 2,000 requests, every feature, no card. The free requests are given once per account, not every month. Pro has every feature and 50,000 requests a month, and request packs add more when you need them. Team gives each seat everything in Pro, with one pool of requests, one Cage policy for the org and one bill. Your own Claude or OpenAI account pays the model, and we add no fee on tokens. See Plans and limits. For Team or Enterprise, write to us on Discord.

Is there a free plan?

Yes. Start free: 2,000 requests, every feature, no card. The 2,000 requests are given once per account; they do not come back each month. When they run out, nothing breaks: your requests go straight to the model until you move to Pro. See what Free includes.

What happens when I run out?

Nothing breaks. Your requests go straight to the model, so your agent keeps working. Context Saving, Memory, Skills and Thinking in Code pause; Cage core protections stay on. A session that is already running keeps every feature until it ends, within a limit. You are told once, not on every turn. To resume, move to Pro, buy a request pack, or wait for your next month on Pro. See what pauses and what stays on.

Will it lower my bill?

It depends on your work. In paired runs on Claude Opus 5.5, cost was 25 to 29% lower on agentic coding, 22% lower on test-driven debugging, 21 to 22% lower on large tool output and 15% lower on multi-agent orchestration. Task success was equal in both arms. We do not quote a figure for your account. The console measures your saving from your own traffic. See Benchmarks.

Do Memory and Skills cost tokens?

Yes. On Claude Haiku 4.5, the memory block added 534 to 548 uncached input tokens to a new session's first request once earlier sessions had saved facts. The paired runs of 2026-09-29 ran with memory and skills on, so their results include that cost.

Do you hold my Claude or OpenAI credential?

No. The gateway forwards it with each request. Your provider bills you directly, and we do not resell tokens.

How does Context Mode compare with other tools?

See the landscape: gateways, token optimizers, agent firewalls, memory layers and more.

Data

Does the gateway see my prompts?

Yes. It sits on the request path, so it reads and rewrites every request your agent sends. There is no setup in which your prompts do not pass through it.

What do you store?

How long do you keep it?

DataKept for
Operations rows: timing, tokens, cost and errors per request, read by us to debug30 days
Savings ledger, by day90 days
Cage activity30 days by default. You can set 1 to 365.
PromptsNo automatic expiry yet. Ask us and we delete them.

Where is it stored?

On Cloudflare. Each account's archive and index are separate Durable Object instances.

Can I delete or export my data?

Not yet from the console. There is no one-click delete or export today. Ask us and we delete your account's data. You can export your Cage policy.

Control

Can I turn parts off?

Memory and Skills are on by default. Switch either off in the console, under Settings, in "What your agent gets". Each skill also has its own switch on the Skills screen. Memory and Skills explain what each one adds.

How do I leave?

Remove ANTHROPIC_BASE_URL and ANTHROPIC_CUSTOM_HEADERS from ~/.claude/settings.json. In ~/.codex/config.toml, remove model_provider = "context-mode". The CLI left a backup of each file next to it. Your agent then talks to the model API directly.

Which models does it work with?

The ones Claude Code uses through Anthropic, and the ones Codex uses through OpenAI with a ChatGPT login or an API key. Our cost figures are measured on Claude Opus 5.5 and Claude Haiku 4.5, with Claude Code. We have not measured Codex cost yet.