Context Mode vs LLM gateways, token optimizers and agent firewalls

Choose Context Mode Gateway when you run Claude Code or Codex and want lower cost, one tool-call policy and one memory for both. Choose a general LLM gateway when you need many model providers, self-hosting or SSO, which we do not ship yet.

Context Mode Gateway sits on the request path of Claude Code and Codex. This page puts it next to 97 table entries in ten groups (one vendor can fill more than one row), with Context Mode in the first row of every table: what each tool does, where it runs, and where we win. We read the sources on 2026-09-29 and 2026-09-30; the list at the end dates each one. "Not documented" means we did not find it on the pages we read.

Key facts

What it is
A dated comparison of Context Mode Gateway with LLM gateways, token optimizers, agent firewalls, sandboxes, memory layers, observability tools, native client features and skills.
Who it is for
CTOs and engineers choosing tools for Claude Code or Codex.
Clients
Claude Code and Codex. Anthropic and OpenAI models, paid by your own Claude or OpenAI account.
Measured result
Not measured on its own. For the gateway as a whole, see Benchmarks.
Limits
Sources were read on 2026-09-29 and 2026-09-30, and tools change. "Not documented" means we did not find it on the pages we read.

Choose Context Mode when …

  • You run Claude Code and Codex and want one set of controls for both. One Cage policy, one memory store, one archive and one skill catalog reach both clients on every machine. Setup is one config change per agent, which the Context Mode CLI writes; nothing keeps running on the laptop. We checked 28 gateway mechanisms against Codex, most with a Codex wire test. Open: tool-schema trimming (at most about 3k cached tokens a request), and part of the Codex saving is not counted yet. See Quick start.
  • You want a saving you can check. Claude Code on Opus 5.5 cost 65.5% less through the gateway on a long agentic coding session, which went from $7.20 to $2.49, 4 of 4 pairs cheaper. Over 168 shorter paired runs, 155 cost less and 168 of 168 tasks passed in both arms. Every pair of the rows shown is published on Benchmarks. A reply-style block asking for short answers was on in every Context Mode run. That is cash on API billing; on a Claude or ChatGPT plan it is list-price value, not cash.
  • You want one tool-call policy on every machine, with no hook to install. Cage checks each call before it runs, on the gateway, for Claude Code and Codex on every machine pointed at it. It reads a call by what it does: 22 capabilities, through 98 wrapper words and 33 carriers such as sudo, sh -c, xargs and ssh. Every block goes to one log you export as CSV or JSON Lines. CC Safety Net and dcg also parse what a command does, with a hook and a log on each machine.
  • You want memory that drops stale facts. A changed fact closes the old row and keeps it as history. In a controlled test on 24 facts with the gateway's own code, 0 of 24 memory blocks held an old value, against 23 of 24 for an append-only store, and the store answered 38 of 38 "what was true then" questions. See Memory.
  • You want your own model account to keep paying, with no fee on tokens. Your Claude or OpenAI account pays the model, and we add no fee on tokens. OpenRouter lists a 5.5% platform fee on its Standard plan, and Requesty bills "what your application spends, plus 5%" (OpenRouter pricing, Requesty pricing, read 2026-09-30).

Choose another tool when you need many model providers, your own cloud, SSO and team roles, or nothing leaving the machine. Each group below names which one, and Limits today lists ours.

Summary

The nine Gateway features, against the seven tools this page names most. Each cell comes from the group tables below and their sources. "Not documented" means we did not find it on the pages we read.

Which plan has what. Every plan has every feature: Context Saving, Thinking in Code, Search, Session resume, Context limit, Reply, Live (each request as it happens, in the console), Analytics (tokens, cost and saving over time), Cage, Memory and Skills. Free gives 2,000 requests once, with no card. Pro gives 50,000 requests a month, and request packs add more. Team pools the requests of its seats, with one Cage policy for the org and one bill. When the requests run out, nothing breaks: requests go straight to the model. See Plans and limits.

FeatureContext ModeClaude CodeCodexCursorHeadroomRTKLiteLLMPortkey
Context SavingFolds tool output once, on the request path; the raw output is archivedPrompt caching, auto-compaction, a cap on MCP outputAutomatic compaction, a limit on tool output tokensNot documentedCompresses tool output, logs and history on your machine; keeps the originalsFilters output of 100+ shell commands on each machine; saves the full output when a command failsNot documentedCaching
Thinking in CodeThe agent's script runs in a Worker isolate; what it prints reaches the modelNot in the client; Anthropic's programmatic tool calling is for apps you buildThe unified exec tool runs commands on your machine; code mode is off by defaultNot documentedNot documentedNot documentedNot documentedNot documented
CageChecks each call before it runs, by what it does; an exportable logPermissions, hooks, managed settings; can ask a personSandbox, approvals, requirements.toml; can ask a personRun modes and a sandboxNot documentedRTK Pro: org-wide policies that block dangerous commandsAllows or denies tool calls by name and argument valuesMCP guardrails on tool-call arguments; request guardrails check the last message
MemoryOne store for both clients, on every machine; old facts closed, not deletedCLAUDE.md and auto memory, on one machineMemories, on one machine, off by defaultNot documentedOne store across agents, on one machineNot documentedNot documentedNot documented
SkillsOne catalog on the gateway; at most one skill a turn, named in the replySKILL.md on disk; managed settings deploy themSkills folders and an admin folderRules; team rules can be requiredNot documentedRTK Pro shared skillsA Claude Code plugin marketplaceSyncs team skills to SKILL.md files
SearchPrompts, commands and archived tool output, for the agentNot documentedNot documentedNot documentedFetches originals back on demandNot documentedRequest logs for peopleLogs for people
Session resumeOne command rebuilds a session in either client, on any machine--resume from the local file; --teleport for web sessionscodex resume from the local fileNot documentedNot documentedNot documentedNot documentedNot documented
Context limitKeeps each thread under your limit; moved turns are archivedAuto-compaction; the local session file keeps the full transcript, and the client gives the model no search over itAutomatic compactionNot documentedCompresses historyNot documentedFalls back to a model with a larger windowNot documented
ReplyOne reply style for both clientsOutput stylesA line in AGENTS.md can ask for a styleA rule can ask for a styleAlso trims output tokensNot documentedNot documentedNot documented
ClientsClaude Code and Codex, one policyClaude CodeCodexCursorClaude Code, Codex, Cursor, Copilot and othersClaude Code, Codex and others; 18 tools per its READMEAny client of the OpenAI or Anthropic API formatClaude Code, Codex
Runs whereA hosted gateway; one config change per agent, which the CLI writesYour machineYour machineYour machineYour machineYour machine; RTK Pro also in your cloudYour own infrastructureHosted, or its open-source gateway
Price, read 2026-09-30Free to start, 2,000 requests with no card; then Pro, Team or Enterprise. Plans and limits. Your account pays the modelPart of Claude Pro, $20 a month when paid monthlyCould not fetch (HTTP 403)Individual $20 a month; Teams $40 per user a monthApache 2.0; team option: could not fetchApache 2.0; RTK Pro: Dev $6 a month for 1 seat, Teams from $100 a month; 7-day free trialMIT core; Enterprise: talk to salesDeveloper free; Production $49 a month, plus $9 per extra 100k requests

Three products

Context Mode Gateway in the same terms

runs whereA hosted gateway on the request path, at gateway.context-mode.com, on Cloudflare. Setup edits one config file per agent: the base URL, and for Claude Code ENABLE_TOOL_SEARCH=true, which keeps its tool search on behind a gateway. Nothing keeps running on the machine.
agentsClaude Code and Codex, checked against each other on 28 gateway mechanisms. Other agents that go through the gateway are recorded, not enforced.
license and priceCommercial and hosted. Free to start, then Pro, Team or Enterprise; see which plan has what. Your Claude or OpenAI account pays the model; we add no fee on tokens. Cage is off until you turn it on.
changes the requestYes. It folds tool output, trims tool schemas, adds memory, skills and a reply style, keeps each thread under your context limit, and checks each tool call against Cage.
recordsTokens, cost and saving for each request; on Codex, part of the saving is not counted yet. Every Cage block, and a policy history where each entry's hash covers the one before it.
archiveTool output the gateway shortens, and old turns it moves out, are kept first. The agent reads them back byte for byte, images aside, except secrets and e-mail addresses, which are masked before storage.
your promptsEvery prompt passes the gateway. Secrets and e-mail addresses are masked before the archive, and the Cage log holds no command text. See what we store.
does notRun your code in an OS sandbox. Run in your own cloud. Route to providers beyond Anthropic and OpenAI. Offer SSO, team roles or group scope in Cage yet. Ask a person before a call runs.

Market map

Rows are groups of tools; the first row is Context Mode. Columns are Context Mode features; the last one is Context Mode Insight's job. The Context Mode column says how we answer each group. Each other cell names tools in that group that do some of that job. "None found" means we found no tool in the group that documents it.

GroupContext ModeContext SavingThinking in CodeCageMemorySkillsSearchSession resumeContext limitReplyTeam analytics (Insight)
Context Mode GatewayAll nine features on one hopFolds tool output once; the raw output is archivedScripts run in a Worker isolate; what they print returns22 capabilities, 98 wrapper words and 33 carriers, checked before the call runsOne store per account; old facts closedOne catalog; at most one skill a turn, named in the replyPrompts, commands and archived tool outputOne command, either client, any machineKeeps each thread under your limit; moved turns archivedOne reply style for both clientsInsight: 222 patterns, 13 MCP tools, scoped by role
LLM gatewaysFolds, checks, remembers and measures on one hop, for Claude Code and Codex, with no fee on tokensOpenRouter context compression, Kong prompt compressor (both shorten the prompt)None foundLiteLLM tool permission guardrail; Portkey MCP and request guardrails; Cloudflare and Kong guardrails on message textNone foundLiteLLM plugin marketplace, Portkey skill syncRequest logs for people: Helicone, LiteLLM, Vercel, BifrostNone foundLiteLLM falls back to a larger window; OpenRouter context compressionNone foundBudgets and logs: LiteLLM, Portkey, Helicone, Cloudflare, Vercel
Token and context optimizersFolds on the request path for every machine, and archives the raw outputRTK, Headroom, Caveman, Compresr, LLMLingua; our free Context Mode EngineContext Mode EngineRTK Pro command policiesHeadroom cross-agent memoryRTK Pro shared skillsHeadroom fetches originals back; Context Mode Engine searches its storeContext Mode Engine restores a snapshot after a resumeCompresr summarizes history ahead of compactionCaveman short replies, Headroom verbosity steeringRTK Pro, rtk gain, headroom savings
Web and docs retrievalA 5-minute cache mark on each WebFetch page; no search or crawlFirecrawl, Exa, Tavily, Jina Reader and Context7 return shorter page or docs textNone foundNone foundNone foundContext7 installs a skillNone foundNone foundNone foundNone foundNone found
Agent security and MCP gatewaysCage checks each call before it runs, by what it doesNone foundNone foundCC Safety Net, dcg, Zenity, Lasso, Knostic, Prisma AIRS, Operant, MintMCP, Obot, Docker MCP GatewayNone foundSnyk Agent Scan and Cisco mcp-scanner scan skillsNone foundNone foundNone foundNone foundAudit trails: Zenity, Lasso, MintMCP, Obot
Sandboxes and code executionThinking in Code runs scripts in a Worker isolate; Cage checks the hosts they callNone foundCloudflare Code Mode, for agents you build; E2B, Daytona, Modal and Docker Sandboxes run codeClaude Code /sandbox, Codex sandbox, Cursor sandbox, Docker SandboxesNone foundNone foundNone foundNone foundNone foundNone foundNone found
Memory layersOne store for both clients; old facts closed, not deletedNone foundNone foundNone foundmem0, Zep and Graphiti, supermemory, claude-mem, Basic Memory, Letta Code, Cognee, HeadroomLetta Codeclaude-mem, Basic MemoryNone foundNone foundNone foundNone found
Session history and resumeOne copy on the gateway; resume in either client; search the archiveNone foundNone foundNone foundNone foundSpecStory Lore turns history into skillsSpecStory Cloud, Entire, cass, Claude Code History ViewerEntire, SpecStory, cassNone foundNone foundSpecStory Cloud analytics
Observability and team analyticsTokens, cost and saving per request, read from the whole requestDatadog finds waste and suggests a fixNone foundNone foundNone foundNone foundTraces for people: Langfuse, LangSmith, Phoenix, BraintrustNone foundNone foundNone foundLangfuse, Datadog, Jellyfish, LinearB, DX, Swarmia, Faros AI, ccusage
Native client featuresOne set of features across both clients and every machineClaude Code and Codex compaction, Claude Code tool search, Anthropic context editing, prompt cachingAnthropic programmatic tool calling, Codex execClaude Code permissions and hooks, Codex approvals and requirements.toml, Cursor run modesClaude Code auto memory, Codex memories, Anthropic memory toolClaude Code and Codex skillsNone foundclaude --resume, codex resume, --teleport from Claude Code on the webClaude Code and Codex auto-compaction, Anthropic compaction and context editingClaude Code output stylesClaude Code analytics and OpenTelemetry, Cursor analytics
Skills and rulesOne catalog on the gateway; at most one skill a turn, named in the replyNone foundNone foundRules files ask; Cursor can make team rules requiredCLAUDE.md, AGENTS.mdAgent Skills, skills.sh, anthropics/skills, Cursor, Cline, Roo and Continue rulesNone foundNone foundNone foundA line in CLAUDE.md or AGENTS.md can ask for a styleNone found

Four questions for each tool

LLM gateways

A gateway sits between the agent and the model provider. Most route requests to many providers, set budgets and keep logs.

ToolRuns whereAgents namedLicenseWhat it doesWhat it does not do
Context Mode GatewayHosted on Cloudflare; one config change per agentClaude Code and Codex, one policyCommercial; free to start, then Pro, Team or Enterprise. Your Claude or OpenAI account pays the model, with no fee on tokensFolds tool output once and archives the raw output, trims tool schemas, keeps each thread under a context limit, checks each tool call with Cage, adds memory, skills and a reply style. 65.5% lower cost on a long Claude Code session.Routes to Anthropic and OpenAI, not to other providers. Hosted only. No SSO or team roles yet.
LiteLLMSelf-hosted proxy: Docker, Kubernetes, AWS or GCPClaude Code, Codex CLI, Claude Desktop, any client of the OpenAI or Anthropic API formatMIT core; its enterprise folder is under a commercial licenseRoutes to 100+ providers. Virtual keys, budgets and logs. A tool permission guardrail allows or denies tool calls by name and checks argument values; with on_disallowed_action, Block halts the request and Rewrite strips the forbidden tools and returns an error message inside the response. An MCP gateway and a Claude Code plugin marketplace. Falls back to a model with a larger window when the context is full.Folding tool output, memory and session resume are not documented.
Portkey, now Prisma AIRS AI GatewayHosted; its open-source gateway also runs on Node, Docker, Workers or KubernetesClaude Code, CodexGateway MIT; platform commercial. Palo Alto Networks bought Portkey in May 2026.1,600+ models, budgets, fallbacks, caching, RBAC, an MCP gateway. MCP guardrails check tool-call arguments and results; a Request Parameters Check allows or blocks the tools a request offers. Syncs team skills to SKILL.md files.Its guardrails on the model request check the last message. Checks on built-in shell tool calls, folding, memory and resume are not documented.
HeliconeCloud, or self-hosted with Docker or HelmClaude CodeApache 2.0. Mintlify bought Helicone. Its post of Mar 3, 2026 says its services stay live in maintenance mode, and security updates, new models, and bug and performance fixes all keep shipping.Request logs, cost, sessions, caching, rate limits, prompts.Its sessions group requests for viewing; they do not resume. Its Claude Code setup gets no new features. Tool-call policy is not documented.
OpenRouterCloudClaude Code (a setup guide)CommercialFailover between providers, org budgets, analytics. Its context-compression plugin removes or cuts messages from the middle of the prompt until it fits the window; it is on by default for endpoints with 8k (8,192 tokens) of context or less.Tool-call policy, an archive of what it drops, and memory are not documented.
Cloudflare AI GatewayCloudflare's networkClaude Code, through Anthropic, Bedrock or VertexCommercial, on all plansAnalytics, logs, cache, rate limits, retries and fallbacks, BYOK. Guardrails check prompts and responses for unsafe content, then flag or block.Checks on tool calls and folding are not documented.
Kong AI GatewayKonnect control plane; data plane self-hosted or on KubernetesClaude Code, Codex CLI, Qwen CodeCommercialRouting, semantic cache, token rate limits, prompt guard, PII redaction, MCP and A2A. A prompt compressor, on the Enterprise tier, uses LLMLingua 2 and drops text.Memory, resume and tool-call rules that read shell commands are not documented.
Vercel AI GatewayVercel's cloudAbout 30 agents through one setup command; Claude Code and Codex have their own endpointsCommercial; no markup on token pricesRouting, fallbacks, budgets, request logs, BYOK. Works with a Claude subscription.Folding, tool-call policy, memory and resume across machines are not documented.
Claude apps gatewayYour own cloud: Kubernetes, Cloud Run, ECSClaude Code and Claude DesktopFirst-party (Anthropic); built into the claude binarySSO by identity-provider group, model allowlists, managed settings for each group (permission rules, hooks, an egress allowlist), spend limits, OpenTelemetry, audit events, failover across Bedrock, Vertex, Foundry and the Anthropic API.No Codex. Its policy runs in the client through managed settings. Rewriting requests is not documented.
BifrostLocal or self-hosted. Its Edge version, in alpha, runs on each company machine.Claude Code, Codex CLI, Gemini CLI, Cursor, Copilot and othersApache 2.0, plus an enterprise edition23+ providers, virtual keys, budgets, semantic cache, an MCP gateway, guardrails, audit logs.Tool-call policy and folding are not documented.
agentgatewayStandalone or on KubernetesClient guides for Claude Code, Claude Desktop, Codex, Cursor, Copilot and othersApache 2.0 (Linux Foundation)A proxy for LLM, MCP and A2A traffic, with RBAC and a CEL policy engine.Folding, memory and shell-command policy are not documented.
claude-code-routerYour machineClaude Code, Codex, Grok CLI, Kimi CLI, OpenCode and othersMITSends each agent to the models and providers you pick, with fallbacks and request logs.Compression, tool-call policy and team policy are not documented.
RequestyCloud, with EU routingClaude Code (a setup guide)A hosted service; license not documentedRouting, fallbacks, load balancing, caching, spend limits, guardrails, SSO, RBAC and an MCP gateway.Folding, memory and shell-command policy are not documented.
TrueFoundry AI GatewayYour VPC, on-prem, air-gapped or across cloudsNot documented on the page we readCommercial1,600+ models, policy control, real-time monitoring and failover, SSO, RBAC and audit logs, plus an MCP gateway, an agent gateway and an agent skills registry.Coding-agent setup, folding and shell-command policy are not documented on the page we read.
Azure API Management AI gatewayAzureNot documented on the page we readCommercialToken limits and quotas per API consumer, semantic caching, content safety checks, remote MCP servers and A2A APIs.Folding and tool-call policy are not documented.
Agent Router, formerly Envoy AI GatewaySelf-hosted on Envoy, or as a service hosted by TetrateAny OpenAI-compatible clientApache 2.0 (Agentic AI Foundation)One API for models and one router for MCP tools, with credentials, limits and failover.Folding, memory and tool-call policy are not documented.

Left out: TensorZero, whose repository was archived on 2026-06-12, and Martian and Unify, which now sell other products.

Where Context Mode differs

  • It changes what the request carries. Most of these gateways route, cache and log; OpenRouter and Kong can shorten a prompt to fit. Context Mode folds tool output the first time the model sees it and sends it the same way on later turns, so the prompt cache keeps reading it. It trims tool schemas and archives the raw output. See Context Saving.
  • Tool calls are read as shell commands. LiteLLM's tool permission guardrail matches tool names and argument values, and Portkey's MCP guardrails check the arguments of MCP tool calls. Cage reads a shell command through about 130 wrappers and carriers, such as sudo, sh -c, xargs and ssh. See Cage.
  • One set of features for two clients. Many gateways route both Claude Code and Codex. Context Mode applies folding, Cage, memory, skills, the reply style and the context limit to both.
  • A saving you can check. On a long session on Opus 5.5, 65.5% lower cost than Claude Code going direct, from $7.20 to $2.49. 155 of 168 shorter runs were cheaper. See Benchmarks.
  • No fee on tokens. Your own account pays the model. OpenRouter lists a 5.5% platform fee on Standard and 8% on Business, Requesty bills spend "plus 5%", and Cloudflare adds 5% to credits bought through Unified Billing (pricing pages, read 2026-09-30).

When to pick another tool

Choose LiteLLM, Portkey, OpenRouter or the Claude apps gateway instead when you need many model providers, your own cloud, or SSO and roles today.

Token and context optimizers

These tools cut what the agent reads. Most run on each developer's machine.

ToolRuns whereAgents namedLicenseWhat it doesWhat it does not do
Context Mode GatewayThe request path, for every machine pointed at the gatewayClaude Code and CodexCommercial; in every planFolds tool output the first time it passes, so the cache keeps reading it, and archives the raw output. The agent reads it back byte for byte, images aside, with secrets and e-mail addresses masked. Paired runs on Opus 5.5 with every feature on: 65.5% lower cost on a long session. We have not yet split out the share that folding makes.Run with no network. Trim tool schemas for Codex yet, worth at most about 3k cached tokens a request.
RTKA binary and an agent hook on each machineClaude Code, Codex, Cursor, Gemini CLI, Windsurf, Cline and others; its README says it supports 18 AI coding toolsApache 2.0Rewrites shell commands to rtk versions that filter, group and cut output, for 100+ commands. rtk gain shows estimated savings.Works on shell output. Its README says token counts are estimates. Saves the full output to a local SQLite store when a command fails or is cut; the agent reads it with rtk recall. Output from successful runs is kept only if you turn that on.
RTK ProA SaaS tenant, your own cloud (Helm) or on-prem"Every agent in your organization"; no listCommercial; Dev $6 a month for 1 seat, Teams from $100 a month for 5 seats, Business from $750 a month for 50 seats, Enterprise custom; a 7-day free trial on Dev and TeamsSecret rewriting, token views by team and project, org-wide policies that block dangerous commands, shared skills, sandboxes.A full-output archive and measured savings are not documented.
HeadroomYour machine, as a library, a local proxy or an MCP server; a paid option for teamsClaude Code, Codex, Cursor, Copilot, Cline, Aider, Goose and othersApache 2.0Compresses tool output, logs, files and history. Keeps the originals on your machine and fetches them on demand. Leaves the cached prefix unchanged and compresses only new bytes. Also trims output tokens and keeps a memory store shared across agents. Its README reports 21 to 57% fewer tokens on four scenarios built from MCP output, and 90% on repeated JSON and log lines.Tool-call policy and memory shared across machines are not documented. Its savings are token counts, not paired cost runs at list price.
CavemanYour machine: a skill, a local proxy, middleware30+ agents, among them Claude Code, Codex, Cursor and GeminiSkill and CLI MIT; proxy BSL 1.1Makes the agent answer in short text. The proxy shrinks what the agent reads.Its README says no reviewed API benchmark is published yet.
Compresr Context GatewayA local proxyClaude Code, Cursor, OpenClawApache 2.0Builds the history summary in the background, so compaction is instant.Measured savings, policy and recovery of dropped text are not documented.
LLMLinguaA library in your codeNone directly; works with LangChain and LlamaIndexMIT (Microsoft)Compresses prompts by dropping tokens the model can do without.A coding-agent setup is not documented. The text it drops is gone.
Augment Context EngineAn MCP server, local or hostedClaude Code, Codex, Cursor, Gemini CLI, Copilot and othersCommercialCode search by meaning across repos, history and docs, so the agent reads less. Reports 32% fewer tokens on its own benchmark.Does not change the rest of the request. Its figure is a benchmark comparison, not a paired cost run.
Claude ContextAn MCP server, with Milvus or Zilliz Cloud and an embedding APIClaude Code, Codex CLI and other MCP clientsMITCode search by meaning. Its README reports about 40% fewer tokens for equal retrieval quality.Needs a vector database and an embedding key. Retrieval only.
SerenaAn MCP server on your machine, using language serversAny MCP clientOpen source; a paid JetBrains pluginFinds and edits code by symbol instead of reading whole files.A measured cost figure is not documented.
RepomixA CLI, a website or an MCP serverAny LLMOpen sourcePacks a repository into one file for a model, with an optional compressed form and token counts per file.Works on repository input only; tool output and history are out of scope.

Where Context Mode differs

  • On the request path. Every request through the gateway is folded, from Claude Code or Codex, on any machine that points at it. Nothing keeps running on the machine.
  • Folded once, so the cache holds. A result is folded the first time it passes. Anthropic's docs say that clearing old tool results later costs cache writes. Headroom also leaves the cached prefix unchanged.
  • Folded output is archived. The raw output goes to your archive before it is shortened, and each cut says what was dropped. The agent reads it back byte for byte, images aside, with secrets and e-mail addresses masked, from any machine. Headroom and RTK also keep originals, on the machine that made them.
  • Measured in dollars. On a long session on Opus 5.5, 65.5% lower cost than Claude Code going direct, from $7.20 to $2.49, 4 of 4 pairs cheaper. We have not yet split out the share that folding makes. Headroom and Augment publish token counts. See Benchmarks.

When to pick another tool

Choose RTK or Headroom instead when nothing may leave the machine or you use agents beyond Claude Code and Codex; one person on one machine can start with our free Context Mode Engine.

Context Mode Engine among local optimizers

runs whereYour machine: an MCP server plus hooks in the client. Tools without hooks use a routing file such as AGENTS.md.
agents18 AI coding tools, among them Claude Code, Cursor, Copilot, Codex and Gemini CLI.
licenseFree, source-available under the Elastic License 2.0. No account. 24,230 GitHub stars on 2026-09-30.
what it doesRuns large tool output in a sandbox, indexes it in a local SQLite FTS5 store the agent searches, lets the agent write a short script so only its output enters the context, and restores a snapshot after compaction or a resume.
does notEnforce one policy across machines or keep one store across machines. The gateway does those for Claude Code and Codex.

RTK and Headroom are its closest neighbours. RTK rewrites shell commands; Headroom compresses through a local proxy. Context Mode Engine works through MCP tools and hooks, and keeps the full output in a store the agent can search. See Context Mode Engine.

Web and docs retrieval

These tools hand the agent shorter text for a web page or a library's docs, so it reads less.

ToolRuns whereAgents namedLicenseWhat it doesWhat it does not do
Context Mode GatewayThe request pathClaude CodeCommercialMarks each Claude Code WebFetch page for a 5-minute cache, so the next read of that page within 5 minutes is a cache read. The first read costs 1.25 times a fresh read.Search or crawl the web. Use these tools with it.
FirecrawlA hosted API, or self-hostedClaude Code, Antigravity, OpenCode and MCP clientsAGPL-3.0; the cloud has more featuresSearch, scrape and crawl, returned as clean Markdown or structured data.Changing the rest of the request and tool-call policy are out of scope.
ExaA hosted search API and MCP serverCodex, Claude (a plugin) and MCP clientsMCP server MIT; the API is commercialWeb search with highlights, summaries and page crawling.Changing the rest of the request is out of scope.
TavilyA hosted API and MCP server, remote or localClaude Code, Cline and MCP clientsMCP server MIT; the API needs a keySearch, extract, map and crawl tools.Changing the rest of the request is out of scope.
Jina ReaderA hosted API: a URL prefixAny agent that can fetch a URLApache 2.0Turns a URL into LLM-friendly text, and a query into search results.Changing the rest of the request is out of scope.
Context7A hosted service, through a CLI with a skill or an MCP serverCursor, Claude Code and other coding agentsClient MIT; its API, parser and crawler are privateUp-to-date library docs, fetched when the agent needs them.Covers library docs only.

Where Context Mode differs

  • It does not search or crawl. Use these tools with it. The gateway gives each Claude Code WebFetch page a 5-minute cache mark. See the page cache.

When to pick another tool

Choose Firecrawl, Exa, Tavily, Jina Reader or Context7 when the agent must search, crawl or read library docs; they work next to the gateway.

Agent security and MCP gateways

These tools decide what a coding agent may run, reach or read. Some hook into the client, some proxy MCP traffic, and some scan before use.

ToolRuns whereAgents namedLicenseWhat it doesWhat it does not do
Context Mode CageThe gateway, before Claude Code or Codex runs the callClaude Code and Codex, one policyCommercial; in every planReads each call by what it does: 22 capabilities, through 98 wrapper words and 33 carriers. Monitor mode. An API key can only tighten the policy. Log export as CSV or JSON Lines; a hash-chained policy history.Ask a person first. SSO or group scope. See inside a script file or a compiled program; keep your sandbox on.
CC Safety NetA hook on each machineClaude Code, Codex, Cursor, Gemini CLI, GitHub Copilot CLI, OpenCode and othersMITParses what a command does, so a wrapper or reordered flags do not hide it. Blocks SSH keys, .env files and the credential files coding CLIs keep. Rulebooks for Terraform, AWS, gcloud and Azure.Each machine installs the hook; the policy is shared through git. Its audit trail stays on the machine.
dcgA hook on each machineClaude Code, Codex CLI, Gemini CLI, Copilot CLI, Cursor and othersA custom license based on MIT, with an OpenAI/Anthropic rider50+ packs for databases, Kubernetes, Docker, clouds and Terraform. Scans heredocs and inline scripts. An ask rule where the hook protocol supports it.Config lives on each machine. Malformed hook input is allowed with an audit warning unless you opt into fail-closed.
ZenityHooks on each machine, and an MCP gatewayClaude Code, Codex, Copilot, CursorCommercialBlocks inline, one central policy, audit through hooks and OpenTelemetry.Changing requests or cost is not documented.
LassoClient hooks, rolled out with managed settings; scanning in Lasso's cloudClaude Code, Cursor, Codex, OpenCodeCommercial; its Claude Code hook is MITScans content for injected instructions, checks tool calls before they run, flags or blocks, keeps an audit trail.Its open-source hook warns but does not block. Changing requests or cost is not documented.
Knostic KirinAn IDE extensionCopilot, Cursor, Claude Code, WindsurfCommercialBlocks risky parts and unsafe actions, with a central audit.Codex is not named.
Prisma AIRS (Palo Alto Networks)Endpoint, network and cloud; an inline AI gateway from PortkeyCursor, Claude Code, Codex, AntigravityCommercialOne policy across agents, blocks with an exception request, session timelines. The gateway covers LLM, MCP and A2A traffic.Context savings and an archive are not documented.
Prompt Security (SentinelOne)A commercial platform with an MCP gatewayNames ChatGPT, Gemini and ClaudeCommercialVisibility, blocks prompt injection and data leaks. Its MCP gateway covers 13,000+ known MCP servers.Checks on coding-agent tool calls are not documented in the release.
Lakera Guard (Check Point)A screening APIAny appCommercialScreens prompts and outputs for injection, personal data and bad links.Checks on coding-agent tool calls are not documented.
Operant AIAn MCP gateway and endpoint agentClaude Code, CursorCommercialAllows, blocks or redacts in real time.Codex is not named.
MintMCPA hosted MCP gateway and agent monitorClaude Code, Cursor, ChatGPTCommercialRBAC, audit, scanning for personal data and secrets.Covers MCP and tool calls only.
ObotAn MCP gateway, self-hosted or cloudClaude, Cursor, VS CodeMITRBAC, audit of tool calls and LLM requests, finds MCP servers nobody approved.Shell commands are out of scope.
Docker MCP GatewayLocal containersVS Code, Cursor, Claude Desktop, Claude CodeMITRuns each MCP server in its own container and manages its credentials and OAuth.Covers MCP traffic only; shell commands and model requests are out of scope.
Snyk Agent ScanA local CLI; needs a Snyk tokenClaude Code, Codex, Cursor, VS Code, Copilot, Gemini CLI and moreApache 2.0Scores MCP servers and skills for prompt injection, destructive tools and malware.Scans only; it does not block at run time.
Cisco mcp-scannerA local CLI, an API or a gatewayReads the configs of Cursor, Claude, VS Code and WindsurfApache 2.0Scans MCP servers with YARA rules, an LLM and an API.Scans only.
Invariant GatewayA proxy, hosted or self-hostedFrameworks such as Swarm, Autogen, OpenHands and SWE-agentApache 2.0Traces, and guardrails that block or log.Claude Code and Codex are not named. Last change on GitHub: 2025-11-06.

Where Context Mode differs

  • The check is on the request path. Once you turn Cage on, it decides before Claude Code or Codex runs a call. A blocked command never runs; the terminal runs only an echo that carries the refusal. On each machine it depends on one config change, which the CLI writes. A machine without it is outside Cage: its calls never reach the gateway.
  • One policy for both clients, with no install. The same rules cover Claude Code and Codex on every machine pointed at the gateway. CC Safety Net and dcg read commands well too, with a hook on each machine and a log on each machine.
  • The same hop saves and records. The hop that checks the call also folds the output, archives it and measures the saving.
  • An audit trail you can check. Every block goes to a decision log you can export. Each policy change records who, how and why, and an edited or deleted entry breaks the hash chain.

When to pick another tool

Choose CC Safety Net or dcg instead when you want a free check on one machine with no network, and Prisma AIRS, Zenity or Lasso when you need SSO, group policies or models that detect injection.

Sandboxes and code execution

A sandbox limits what a program can reach once it runs: files, network, the rest of the machine. Code execution for agents runs code the model writes, so only the result enters the context.

ToolRuns whereAgents namedLicenseWhat it doesWhat it does not do
Context Mode Thinking in Code and CageCloudflare Worker isolates for the agent's scripts; the gateway for CageClaude Code and Codex, with no code changeCommercialThe agent writes a short script; it runs in a Worker isolate, and what it prints reaches the model. With Cage on, every host the script calls is checked.Isolate your machine. Run long jobs or other languages: scripts are short JavaScript, and local files go through your agent's shell.
Claude Code /sandboxYour machine: Seatbelt on macOS, bubblewrap and seccomp on LinuxClaude Code; its runtime can wrap any processRuntime Apache 2.0File and network limits through a local proxy with a domain allowlist. Admins can require it.Anthropic's page says it is "not a complete isolation boundary". Without its dependencies it runs commands outside the sandbox unless an admin turns that off.
Codex sandboxYour machine: Seatbelt, bubblewrap and seccomp, a Windows sandboxCodexApache 2.0Network off by default. Admins lock sandbox modes, approval policies and domain rules in requirements.toml.Its settings do not limit web search, browser tools, MCP connections or model API calls.
Cursor sandbox and run modesYour machine: Seatbelt on macOS, Landlock and seccomp on LinuxCursorCommercialAuto-review, allowlist and run-everything modes, network rules in sandbox.json, admin overrides from the dashboard.Cursor says its auto-review classifier "can make mistakes". Cursor only.
Docker SandboxesA microVM, local or in the cloudClaude Code, Codex, Copilot, Cursor, Gemini, Kiro, OpenCode and othersCommercial; org policy is on a paid planIsolates the agent in a microVM. A host proxy applies network policy and injects credentials.Docker's page says "the agent has full control inside the VM". Single tool calls are not read.
Claude Code dev containerDocker, local or CodespacesClaude CodeA reference setup in an open-source repoAn isolated environment with a reference firewall script.With permission prompts skipped, Anthropic says it does not stop a malicious project from sending out what the container can read.
E2BCloud Linux VMs, or self-hosted with TerraformAny agent, through its SDKApache 2.0Starts a VM on demand to run agent code.Policy on the tool calls of an agent on a laptop is not documented.
DaytonaCloudAny agent, through its SDKAGPL-3.0 (LICENSE at v0.190.0); core development private since June 2026Sandboxes with their own kernel, file system and network stack.Policy on the tool calls of an agent on a laptop is not documented.
Modal SandboxesModal's cloud, on gVisorAny agent, through its SDKCommercialRuns code with the network blocked or limited to CIDR ranges or domains.A sandbox lives 5 minutes by default and 24 hours at most. No tool-call policy.
Cloudflare Code ModeCloudflare Workers: the model's code runs in a sandbox cut off from the Internet, reaching only the MCP servers it is givenAgents you build with the Cloudflare Agents SDKPart of the Agents SDKTurns MCP tools into a TypeScript API. The model writes code against it, and only what the code logs with console.log goes back to the agent. Published 2025-09-26.For agents you build. Setup for Claude Code or Codex is not documented in the post.

Where Context Mode differs

  • Thinking in Code for the agents you already use. The agent writes a short script, the gateway runs it in a Cloudflare Worker isolate, and only what it prints reaches the model. Cloudflare Code Mode uses the same pattern for agents you build with the Agents SDK. The gateway gives the pattern to Claude Code with no code change. Codex has its own exec, which runs a JavaScript program on your machine; through the gateway, the script runs off the machine in a Worker isolate. When Cage is on, every host the script calls is checked against it. See Thinking in Code.
  • Cage decides before the call runs. A sandbox limits a program once it runs. Cage stops the call from being run, and the block is in the decision log.
  • One policy for two clients. The client sandboxes are set up per client and per machine.

When to pick another tool

Keep your Claude Code, Codex or Cursor sandbox on in every case, and choose E2B, Daytona, Modal or Docker Sandboxes when a job needs a full Linux machine: a sandbox limits a program once it runs, and Cage reads the call before it runs.

Memory layers

These tools carry what an agent learned into later sessions. Some write files on your machine, some are APIs you call from your own code.

ToolRuns whereAgents namedLicenseWhat it doesWhat it does not do
Context Mode MemoryThe gateway: one store per account, nothing on each machineClaude Code and Codex share one storeCommercial; in every planPicks facts from turns and tools, and adds them when a conversation starts or after compaction. A changed fact closes the old row. In a controlled test on 24 facts: 0 of 24 blocks held an old value, against 23 of 24 for an append-only store; 38 of 38 "what was true then" answers.A single "as of" query yet. Team memory apart from personal memory.
Claude Code memoryLocal files: CLAUDE.md you write, and auto memory notes under ~/.claude/projects with a MEMORY.md index loaded each sessionClaude CodePart of the clientInstructions you write, plus notes Claude writes as it works.Stays on one machine. Codex does not read the auto memory notes.
Codex memoriesLocal files under ~/.codex/memories/CodexPart of the clientSummaries and entries from earlier chats. Off by default.OpenAI's page says to keep required team guidance in AGENTS.md and treat memories as a recall layer. Stays on one machine.
Anthropic memory toolYour application: Claude asks for file operations under /memories, and your code runs them against storage you pickApps built on the Claude APIPart of the APIA building block for your own agent's memory.You write and host the storage. Not a feature of Claude Code or Codex.
mem0 and OpenMemoryA library, a self-hosted server, or mem0's cloudOpenMemory: Cursor, VS Code, Claude and MCP clientsApache 2.0; cloud is commercialCaptures coding preferences and patterns as you work and adds the ones that match the current project.Tool-call policy and output folding are not documented.
Graphiti and ZepGraphiti in your own graph database; Zep as a managed serviceAny MCP client, through its MCP serverGraphiti Apache 2.0; Zep commercialA graph of facts over time. A changed fact is marked invalid, not deleted.Needs an LLM and a graph database. A coding-agent plugin is not documented.
supermemoryCloud, or self-hostedClaude Code, Codex, Cursor, OpenCode and others, by plugin or MCPMITHooks save conversations and important tool use. Extracts facts, handles changes and contradictions, and keeps team memory apart from personal memory.Tool-call policy and output folding are not documented.
claude-memHooks and a local worker with SQLite and Chroma; a hosted tierBuilt for Claude Code; installers for OpenCode, Antigravity and othersApache 2.0Saves observations from each session, compresses them, and searches them through MCP tools.Tool-call policy and output folding are not documented.
Basic MemoryMarkdown files and a local SQLite index; a hosted versionClaude, Codex, Cursor, ChatGPT and MCP clientsAGPL 3.0Notes you and the agent both read and write, with search and links between them.Tool-call policy and output folding are not documented.
Letta CodeIts own agent: a CLI, a desktop app, the browserLetta Code agentsApache 2.0A git-backed memory file system. Agents rewrite their own memory and skills over time.A separate agent, not a layer for Claude Code or Codex.
CogneeLocal, self-hosted, or Cognee CloudClaude Code and Codex, through plugins; MCP clientsApache 2.0; the cloud is commercialTurns documents, code and conversations into a knowledge graph the agent searches. Runs locally on small models with no API key.Tool-call policy and output folding are not documented.
Headroom memoryYour machine, as part of HeadroomClaude, Codex, Gemini and GrokApache 2.0One shared memory store across agents, with automatic dedup. headroom learn writes corrections into CLAUDE.md or AGENTS.md.A store shared across machines is not documented.

Where Context Mode differs

  • Added on the request path. No plugin, hook or MCP server on each machine. The gateway adds memory when a conversation starts or after compaction.
  • One store per account, on every machine. Claude Code and Codex share it. Headroom also shares one store across agents, on the machine that runs it.
  • Old facts are closed, not deleted. A replaced fact stays as history. Graphiti does this too.
  • Measured against the alternatives. In a controlled test on 24 facts with the gateway's own code, closing old rows kept old values out of 24 of 24 memory blocks, where an append-only store let them into 23; an overwriting store could not answer 38 "what was true then" questions. Memory gives the live check, cost included.

When to pick another tool

Choose Claude Code memory or Codex memories instead when you use one client on one machine, and Graphiti, mem0 or supermemory when you need self-hosting, team memory or an API for your own apps.

Observability and team analytics

These tools record what coding agents do. Tracing tools record it through hooks or OpenTelemetry. Engineering analytics join it to git and tickets.

ToolRuns whereAgents namedLicenseWhat it doesWhat it does not do
Context Mode GatewayThe request path; one record for both clientsClaude Code and CodexCommercial; Analytics is in every planTokens, cost and saving for each request, read from the whole request: system prompt, tool schemas, CLAUDE.md, skills and tool output. Every Cage block.Evals, datasets or prompt management. Read git or tickets. On Codex, part of the saving is not counted yet.
LangfuseCloud or self-hosted. Hooks or OpenTelemetry on each machine, or behind an LLM gateway.Claude Code, Codex, Copilot, Cursor, Kiro, OpenCode, Augment Code, VS CodeMIT, except its ee folders. Now part of ClickHouse.Traces, evals, prompt management, cost per person or repo.Hook tracing does not change requests. Its guide says to treat it "as telemetry, not enforcement", and that its hooks cannot read what CLAUDE.md, skills or auto-loaded context added to the prompt.
LangSmithCloud, hybrid or self-hosted. A Claude Code plugin on each machine.Claude CodeCommercialTraces of messages, tool calls, compaction and subagents. Redacts secrets on your machine before upload. Evals.System prompts are not traced. Does not change requests.
Arize PhoenixSelf-hosted or cloud. A Claude Code plugin with hooks; OpenTelemetry.Claude Code, Claude Agent SDKElastic License 2.0Traces, evals, datasets, experiments, a prompt playground.Does not proxy requests.
BraintrustCloud, or hybrid with your own data plane. A gateway in public preview.Claude CodeCommercial; its proxy code is open sourceLogs, evals, playgrounds. The gateway caches, logs and fails over.Changing tool output or checking tool calls is not documented.
Datadog Agent ConsoleDatadog's cloud, from Claude Code's OpenTelemetry or the Anthropic usage integrationClaude Code, Cursor, GitHub CopilotCommercial; in PreviewSpend, sessions, time to merge. Finds problems such as retry loops and reading the same file again, each with a monthly cost and a suggested fix.Blocking or changing agent requests is not documented.
Jellyfish AI ImpactSaaSCopilot, Cursor, Claude Code, Amazon Q, Gemini Code Assist, Windsurf, CodeRabbit, Devin, JulesCommercialAdoption, spend and delivery for each tool and team, from git and planning tools.Request content is not among its documented sources.
LinearBSaaSCopilot, Cursor, Claude Code, Amazon Q Developer, Amazon KiroCommercialAdoption, acceptance and tokens, next to cycle time and change failure rate.Request content is not among its documented sources.
DXSaaSClaude Code, Cursor, GitHub CopilotCommercialA framework for use, impact and cost. Share of AI code by commit, PR and repo. Unused licenses.Request content is not among its documented sources.
SwarmiaSaaSClaude Code, GitHub Copilot, Cursor, Codex, CodeRabbitCommercialCost for each PR and initiative, cycle time, review agents.Self-hosting is not documented.
Faros AISaaSNot documented on its public pagesCommercialLinks AI coding spend to shipped work, and covers model routing and governance.Its docs need a login; details are not documented publicly.
ccusageA CLI on your machine, reading local session filesClaude Code, Codex, OpenCode, Amp, Droid and othersMITDaily, monthly and per-session token use and cost from local data.Reads one machine. Does not change requests.

Where Context Mode differs

  • It sees the whole request. The gateway reads what the model reads: system prompt, tool schemas, CLAUDE.md, skills and tool output. Langfuse's guide says its hooks cannot read what CLAUDE.md, skills or auto-loaded context added.
  • It changes the request, then records the change. Tracers and dashboards record. Datadog finds waste and suggests a fix. Context Mode removes some of that waste on the request path and shows tokens, cost and saving for each request.
  • One record for two clients. Claude Code and Codex requests go to the same account.

When to pick another tool

Choose Langfuse, LangSmith, Phoenix or Braintrust instead when you need evals and prompt management, and Jellyfish, LinearB, DX, Swarmia or Faros AI when you want AI use joined to git and tickets.

Context Mode Insight among team analytics

runs whereOur cloud. A developer opts in with an Insight key, and Context Mode Engine forwards each session event: tool name, file path, command type, exit code, and the first 200 characters of the event's text, which includes the start of each prompt. Secrets, e-mail addresses and the home folder name are masked first. File contents stay on the machine.
agentsThe coding tools Context Mode Engine supports.
licenseCommercial.
what it doesRuns 222 patterns over those events, such as blockers, retry waste and error spikes, and answers through 13 MCP tools, scoped by role: an owner sees the organisation, a manager their teams, a member their own work.
does notRead file contents or change requests. It sees the first 200 characters of each prompt. Link to tickets is not documented.

Jellyfish, LinearB, DX, Swarmia and Faros AI join AI use to git and planning tools. Insight starts from what happens inside agent sessions. See Context Mode Insight.

Native client features

Claude Code, Codex and Cursor ship their own tools for context, safety, memory and resume. They are free, first-party and improve with each release.

ToolRuns whereAgentsLicenseWhat it doesWhat it does not do
Context Mode GatewayThe gateway, outside the clientClaude Code and Codex, checked on 28 gateway mechanismsCommercial; every plan has every featureOne archive, memory store, skill catalog, reply style, context limit and Cage policy for both clients, on every machine.Ask a person before a call runs. Work without a third party on the request path.
Claude CodeThe client, on your machineClaude CodeCommercialPrompt caching, auto-compaction, MCP tool search, a cap on MCP output, permissions and hooks, managed settings, --resume and --continue, skills, analytics.The local session file keeps the full transcript; the client gives the model no search over it. Local sessions and memory stay on one machine; sessions run in Claude Code on the web can move to the terminal with --teleport. Claude Code only.
CodexThe client, on your machineCodexApache 2.0Compaction and resume, automatic compaction, a limit on tool output tokens, memories, sandbox and approvals, admin policy in requirements.toml, skills, the unified exec tool that runs commands, and a code mode that is off by default.The local session file keeps the history; searching it from the model is not documented. Codex only.
CursorThe editor, on your machineCursorCommercialRun modes and a sandbox, rules with required team rules, team analytics and an Admin API.Cursor only.
Anthropic context managementThe Claude APIApps built on the APIPart of the APIServer-side compaction, clearing old tool results and thinking, a memory tool.For people who build their own agent. Anthropic's docs say clearing tool results costs cache writes.
Anthropic tool search and programmatic tool callingThe Claude APIApps built on the API; Claude Code uses tool searchPart of the APITools load only when needed. Code runs in a container, so intermediate results stay out of the context.The app builder has to wire them in.
Prompt caching (Anthropic, OpenAI)The providerAll clientsPart of the APIPrices repeated input far lower than new input.Does not shrink what a request carries.
Claude Code output stylesThe client: Markdown files for the user, the project or managed settingsClaude CodePart of the clientChanges how Claude Code responds. Built-in styles: Default, Proactive, Concise, Explanatory and Learning, or your own.Claude Code only.

Where Context Mode differs

  • Across clients and machines. One archive, one memory store, one skill catalog, one reply style, one context limit and one Cage policy for Claude Code and Codex, on every machine.
  • Folded output is archived. Tool output the gateway folds, and old turns it moves out under the context limit, are archived and can be read back byte for byte, images aside, except secrets and e-mail addresses, which are masked before storage. The gateway keeps each thread under the client's own compaction point.
  • Cage runs outside the client. On each machine it depends on one config change, which the CLI writes. A machine without it is outside Cage.
  • Thinking in Code off the machine. Anthropic's programmatic tool calling needs an app you build. The gateway gives the pattern to Claude Code. Codex's own exec tool runs on your machine; for Codex, the gateway runs the script off the machine in a Worker isolate.
  • One reply style for both clients. Claude Code output styles cover Claude Code. The reply style is set once and reaches Claude Code and Codex.

When to pick another tool

Choose the built-in features alone when you use one client, want no third party on the request path, or need an ask step before a call runs.

Session history and resume

These tools keep what an agent did in a session, so you can search it or pick it up again.

ToolRuns whereAgents namedLicenseWhat it doesWhat it does not do
Context Mode GatewayThe gateway keeps one copy of each conversationClaude Code and CodexCommercialOne command rebuilds a conversation as a native Claude Code or Codex session on any machine. Search finds prompts, commands and archived tool output.Bring back every tool output: resume returns message text, tool names and a summary. Share sessions with a team.
SpecStoryYour machine, saving to .specstory/history/; SpecStory Cloud with a loginClaude Code, Codex CLI, Cursor, Copilot, Gemini CLI, Antigravity, OpenCode and othersCLI open source (Apache 2.0 repository); the cloud is commercialSaves every session locally. With SpecStory Cloud it searches and shares sessions, and resumes them across agents, computers and team members. The cloud also has analytics.Changing requests and tool-call policy are not documented.
EntireA CLI and git hooks; checkpoints are git refs in your repositoryClaude Code, Codex, Cursor, Antigravity, Pi and moreMITSaves each agent session with the commit it made. entire session resume picks up a stopped session by branch, and checkpoints can be searched.Changing requests and tool-call policy are not documented.
cassA TUI and CLI on your machine; syncs other machines over SSHCodex, Claude Code, Gemini CLI, Cursor, OpenCode, Aider and many moreMIT with an OpenAI and Anthropic riderOne index over local session history from many agents, with resume commands.Changing requests is not documented.
Claude Code History ViewerA desktop app or a headless server, offlineClaude Code, Codex CLI, Gemini CLI, Cursor, Cline, OpenCode and othersMITBrowses, searches and analyzes local conversation history.Reads history; resume and request changes are not documented.
Claude Code on the webAnthropic's cloudClaude CodeCommercial; on Pro, Max and Team plans, and some Enterprise seatsRuns sessions in the cloud and moves them to and from the terminal with --cloud and --teleport.Claude Code only.

Where Context Mode differs

  • One copy, either client. The gateway keeps one copy of each conversation that goes through it. One command rebuilds it as a native Claude Code or Codex session on any machine. See Session resume.
  • Search covers the archive. Search finds prompts, commands and archived tool output, and the agent reads that output back byte for byte, images aside, except secrets and e-mail addresses, which are masked before storage.
  • Nothing to sync. The copy is made on the request path, so no local folder, git ref or SSH sync is needed.

When to pick another tool

Choose built-in resume, SpecStory, Entire or cass instead when you need every tool output of a session back, sessions tied to commits, or agents beyond Claude Code and Codex.

Skills and rules

Two different things get called rules. A rules file is text the model reads. A permission is a check that runs before the tool call. The model can ignore text. It cannot skip a check.

ToolWhere it livesAgentsLicenseEnforced or advisoryTeam-wide
Context Mode SkillsThe gateway; imported from any GitHub repository with SKILL.md filesClaude Code and CodexCommercial; in every planThe gateway picks at most one skill a turn, after a second check, and names it in the reply. Policy is enforced by Cage, not by a rules file.One catalog per account, on every machine. No team roles yet.
Claude Code skills and pluginsSKILL.md on the local disk: personal, project, plugin, or the managed settings folderClaude CodePart of the clientGuidance. The model loads a skill when its description fits.Yes. Managed settings deploy skills and require or restrict plugin marketplaces.
Codex AGENTS.md and skillsAGENTS.md in ~/.codex and each folder from the git root down; skills in .agents/skills and /etc/codex/skillsCodexPart of the clientAGENTS.md files are joined into the prompt, 32 KiB by default.An admin skills folder, and plugins for rollout across an org.
Cursor rules.cursor/rules in git, user rules, team rules in the dashboard, AGENTS.mdCursorCommercialGuidance the model reads.Team rules can be required for every member.
Cline rules.clinerules and a global folder; also reads .cursorrules, .windsurfrules and AGENTS.mdCline in VS Code, JetBrains, a CLI and a desktop appApache 2.0Guidance.Not documented.
Roo Code rules.roo/rules, rules for each mode, AGENTS.mdRoo CodeApache 2.0; repository archived on 2026-05-15Guidance.Not documented.
Continue rules.continue/rulesContinueApache 2.0; repository no longer actively maintainedGuidance, added to the system message.Not documented.
Agent SkillsAn open format for skill folders, first developed by AnthropicClaude Code, Codex, Gemini CLI, OpenCode and othersOpen standardNot applicable.Not applicable.
skills.shA directory by Vercel; installs with npx skills addClaude Code, Cursor, Copilot, Cline, Gemini and othersNot documentedNot applicable.Not applicable.
anthropics/skillsA GitHub repository that is also a Claude Code plugin marketplaceClaude CodeApache 2.0; its document skills are source-availableNot applicable.Not applicable.

Where Context Mode differs

  • Skills live on the gateway, not on disk. One catalog reaches Claude Code and Codex on every machine that goes through the gateway. You import from any GitHub repository that holds SKILL.md files. See Skills.
  • The gateway picks the skill. It compares your message with each skill's name and description, runs a second check, loads at most one skill a turn, and names it in the reply. In Claude Code and Codex, the model decides from the descriptions.
  • Policy is a check, not a rules file. Cage decides on the gateway before Claude Code or Codex runs a call. A line in AGENTS.md or CLAUDE.md is text the model may not follow.
  • The cost is stated. The skill list was about 760 tokens in our benchmark accounts. Skills says what we have not measured.

When to pick another tool

Choose Claude Code managed settings, plugin marketplaces or the Codex admin skills folder instead when you must roll skills out to a whole org today.

Limits today

FAQ

Does Context Mode route to other model providers?

No. Anthropic and OpenAI only.

Can I self-host it?

Not on your own servers. On Enterprise, the gateway can run in your own Cloudflare account. See Team and Enterprise.

Does it have SSO or team roles?

Not yet.

Is it an MCP gateway?

No. It is a gateway for the model API. Cage rules also cover calls to tool servers, so one policy covers shell commands, sites, files and MCP tools.

Sources

Read on 2026-09-29, except the lists marked 2026-09-30, which we read or read again on that day.

Something here out of date? Tell us and we fix it.