# Context Mode vs LLM gateways, token optimizers and agent firewalls

Choose Context Mode Gateway when you run Claude Code or Codex and want lower cost, one tool-call policy and one memory for both. Choose a general LLM gateway when you need many model providers, self-hosting or SSO, which we do not ship yet.

Context Mode Gateway sits on the request path of Claude Code and Codex. This page puts it next to 97 table entries in ten groups (one vendor can fill more than one row), with Context Mode in the first row of every table: what each tool does, where it runs, and where we win. We read the sources on 2026-09-29 and 2026-09-30; the list at the end dates each one. "Not documented" means we did not find it on the pages we read.

## Key facts

What it is

A dated comparison of Context Mode Gateway with LLM gateways, token optimizers, agent firewalls, sandboxes, memory layers, observability tools, native client features and skills.

Who it is for

CTOs and engineers choosing tools for Claude Code or Codex.

Clients

Claude Code and Codex. Anthropic and OpenAI models, paid by your own Claude or OpenAI account.

Measured result

Not measured on its own. For the gateway as a whole, see [Benchmarks](https://context-mode.com/docs/benchmarks#pooled).

Limits

Sources were read on 2026-09-29 and 2026-09-30, and tools change. "Not documented" means we did not find it on the pages we read.

## Choose Context Mode when …

- **You run Claude Code and Codex and want one set of controls for both.** One Cage policy, one memory store, one archive and one skill catalog reach both clients on every machine. Setup is one config change per agent, which the Context Mode CLI writes; nothing keeps running on the laptop. We checked 28 gateway mechanisms against Codex, most with a Codex wire test. Open: tool-schema trimming (at most about 3k cached tokens a request), and part of the Codex saving is not counted yet. See [Quick start](https://context-mode.com/docs/quick-start).
- **You want a saving you can check.** Claude Code on Opus 5.5 cost [65.5% less](https://context-mode.com/benchmarks/long-session/ab.json) through the gateway on a long agentic coding session, which went from [$7.20](https://context-mode.com/benchmarks/long-session/ab.json) to [$2.49](https://context-mode.com/benchmarks/long-session/ab.json), 4 of 4 pairs cheaper. Over 168 shorter paired runs, 155 cost less and 168 of 168 tasks passed in both arms. Every pair of the rows shown is published on [Benchmarks](https://context-mode.com/docs/benchmarks#pooled). A reply-style block asking for short answers was on in every Context Mode run. That is cash on API billing; on a Claude or ChatGPT plan it is list-price value, not cash.
- **You want one tool-call policy on every machine, with no hook to install.** [Cage](https://context-mode.com/docs/cage) checks each call before it runs, on the gateway, for Claude Code and Codex on every machine pointed at it. It reads a call by what it does: 22 capabilities, through 98 wrapper words and 33 carriers such as `sudo`, `sh -c`, `xargs` and `ssh`. Every block goes to one log you export as CSV or JSON Lines. [CC Safety Net](https://github.com/kenryu42/cc-safety-net) and [dcg](https://github.com/Dicklesworthstone/destructive_command_guard) also parse what a command does, with a hook and a log on each machine.
- **You want memory that drops stale facts.** A changed fact closes the old row and keeps it as history. In a controlled test on 24 facts with the gateway's own code, 0 of 24 memory blocks held an old value, against 23 of 24 for an append-only store, and the store answered 38 of 38 "what was true then" questions. See [Memory](https://context-mode.com/docs/memory#evaluation).
- **You want your own model account to keep paying, with no fee on tokens.** Your Claude or OpenAI account pays the model, and we add no fee on tokens. OpenRouter lists a 5.5% platform fee on its Standard plan, and Requesty bills "what your application spends, plus 5%" ([OpenRouter pricing](https://openrouter.ai/pricing), [Requesty pricing](https://www.requesty.ai/pricing), read 2026-09-30).

**Choose another tool when** you need many model providers, your own cloud, SSO and team roles, or nothing leaving the machine. Each group below names which one, and [Limits today](#limits) lists ours.

## Summary

The nine Gateway features, against the seven tools this page names most. Each cell comes from the group tables below and their sources. "Not documented" means we did not find it on the pages we read.

**Which plan has what.** Every plan has every feature: Context Saving, Thinking in Code, Search, Session resume, Context limit, Reply, Live (each request as it happens, in the console), Analytics (tokens, cost and saving over time), Cage, Memory and Skills. Free gives 2,000 requests once, with no card. Pro gives 50,000 requests a month, and request packs add more. Team pools the requests of its seats, with one Cage policy for the org and one bill. When the requests run out, nothing breaks: requests go straight to the model. See [Plans and limits](https://context-mode.com/docs/plans-and-limits).

| Feature | Context Mode | Claude Code | Codex | Cursor | Headroom | RTK | LiteLLM | Portkey |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [Context Saving](https://context-mode.com/docs/how-saving-works) | Folds tool output once, on the request path; the raw output is archived | Prompt caching, auto-compaction, a cap on MCP output | Automatic compaction, a limit on tool output tokens | Not documented | Compresses tool output, logs and history on your machine; keeps the originals | Filters output of 100+ shell commands on each machine; saves the full output when a command fails | Not documented | Caching |
| [Thinking in Code](https://context-mode.com/thinking-in-code) | The agent's script runs in a Worker isolate; what it prints reaches the model | Not in the client; Anthropic's programmatic tool calling is for apps you build | The [unified `exec` tool](https://learn.chatgpt.com/docs/config-file/config-reference) runs commands on your machine; code mode is off by default | Not documented | Not documented | Not documented | Not documented | Not documented |
| [Cage](https://context-mode.com/docs/cage) | Checks each call before it runs, by what it does; an exportable log | Permissions, hooks, managed settings; can ask a person | Sandbox, approvals, requirements.toml; can ask a person | Run modes and a sandbox | Not documented | RTK Pro: org-wide policies that block dangerous commands | Allows or denies tool calls by name and argument values | MCP guardrails on tool-call arguments; request guardrails check the last message |
| [Memory](https://context-mode.com/docs/memory) | One store for both clients, on every machine; old facts closed, not deleted | CLAUDE.md and auto memory, on one machine | Memories, on one machine, off by default | Not documented | One store across agents, on one machine | Not documented | Not documented | Not documented |
| [Skills](https://context-mode.com/docs/skills) | One catalog on the gateway; at most one skill a turn, named in the reply | SKILL.md on disk; managed settings deploy them | Skills folders and an admin folder | Rules; team rules can be required | Not documented | RTK Pro shared skills | A Claude Code plugin marketplace | Syncs team skills to SKILL.md files |
| [Search](https://context-mode.com/docs/search) | Prompts, commands and archived tool output, for the agent | Not documented | Not documented | Not documented | Fetches originals back on demand | Not documented | Request logs for people | Logs for people |
| [Session resume](https://context-mode.com/docs/session-resume) | One command rebuilds a session in either client, on any machine | `--resume` from the local file; `--teleport` for web sessions | `codex resume` from the local file | Not documented | Not documented | Not documented | Not documented | Not documented |
| [Context limit](https://context-mode.com/docs/context-limit) | Keeps each thread under your limit; moved turns are archived | Auto-compaction; the local session file keeps the full transcript, and the client gives the model no search over it | Automatic compaction | Not documented | Compresses history | Not documented | Falls back to a model with a larger window | Not documented |
| [Reply](https://context-mode.com/docs/reply) | One reply style for both clients | Output styles | A line in AGENTS.md can ask for a style | A rule can ask for a style | Also trims output tokens | Not documented | Not documented | Not documented |
| Clients | Claude Code and Codex, one policy | Claude Code | Codex | Cursor | Claude Code, Codex, Cursor, Copilot and others | Claude Code, Codex and others; 18 tools per its README | Any client of the OpenAI or Anthropic API format | Claude Code, Codex |
| Runs where | A hosted gateway; one config change per agent, which the CLI writes | Your machine | Your machine | Your machine | Your machine | Your machine; RTK Pro also in your cloud | Your own infrastructure | Hosted, or its open-source gateway |
| Price, read 2026-09-30 | Free to start, 2,000 requests with no card; then Pro, Team or Enterprise. [Plans and limits](https://context-mode.com/docs/plans-and-limits) . Your account pays the model | Part of [Claude Pro](https://claude.com/pricing) , $20 a month when paid monthly | Could not fetch (HTTP 403) | [Individual](https://cursor.com/pricing) $20 a month; Teams $40 per user a month | Apache 2.0; team option: could not fetch | Apache 2.0; [RTK Pro](https://pro.rtk-ai.app/pricing) : Dev $6 a month for 1 seat, Teams from $100 a month; 7-day free trial | MIT core; [Enterprise](https://www.litellm.ai/enterprise) : talk to sales | [Developer](https://portkey.ai/pricing) free; Production $49 a month, plus $9 per extra 100k requests |

## Three products

- **Context Mode Gateway**, the main product, leads every table on this page.
- **Context Mode Engine**, our free plugin, sits among the [local optimizers](#local).
- **Context Mode Insight**, for engineering teams, sits among [team analytics](#insight).

## Context Mode Gateway in the same terms

runs where A hosted gateway on the request path, at gateway.context-mode.com, on Cloudflare. Setup edits one config file per agent: the base URL, and for Claude Code `ENABLE_TOOL_SEARCH=true`, which keeps its tool search on behind a gateway. Nothing keeps running on the machine.

agents Claude Code and Codex, checked against each other on 28 gateway mechanisms. Other agents that go through the gateway are recorded, not enforced.

license and price Commercial and hosted. Free to start, then Pro, Team or Enterprise; see [which plan has what](#plans). Your Claude or OpenAI account pays the model; we add no fee on tokens. Cage is off until you turn it on.

changes the request Yes. It folds tool output, trims tool schemas, adds memory, skills and a [reply style](https://context-mode.com/docs/reply), keeps each thread under your [context limit](https://context-mode.com/docs/context-limit), and checks each tool call against [Cage](https://context-mode.com/docs/cage).

records Tokens, cost and saving for each request; on Codex, part of the saving is not counted yet. Every Cage block, and a policy history where each entry's hash covers the one before it.

archive Tool output the gateway shortens, and old turns it moves out, are kept first. The agent reads them back byte for byte, images aside, except secrets and e-mail addresses, which are masked before storage.

your prompts Every prompt passes the gateway. Secrets and e-mail addresses are masked before the archive, and the Cage log holds no command text. See [what we store](https://context-mode.com/docs/faq#prompts).

does not Run your code in an OS sandbox. Run in your own cloud. Route to providers beyond Anthropic and OpenAI. Offer SSO, team roles or group scope in Cage yet. Ask a person before a call runs.

## Market map

Rows are groups of tools; the first row is Context Mode. Columns are Context Mode features; the last one is Context Mode Insight's job. The Context Mode column says how we answer each group. Each other cell names tools in that group that do some of that job. "None found" means we found no tool in the group that documents it.

| Group | Context Mode | Context Saving | Thinking in Code | Cage | Memory | Skills | Search | Session resume | Context limit | Reply | Team analytics (Insight) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [Context Mode Gateway](https://context-mode.com/docs) | All nine features on one hop | Folds tool output once; the raw output is archived | Scripts run in a Worker isolate; what they print returns | 22 capabilities, 98 wrapper words and 33 carriers, checked before the call runs | One store per account; old facts closed | One catalog; at most one skill a turn, named in the reply | Prompts, commands and archived tool output | One command, either client, any machine | Keeps each thread under your limit; moved turns archived | One reply style for both clients | Insight: 222 patterns, 13 MCP tools, scoped by role |
| [LLM gateways](#gateways) | Folds, checks, remembers and measures on one hop, for Claude Code and Codex, with no fee on tokens | OpenRouter context compression, Kong prompt compressor (both shorten the prompt) | None found | LiteLLM tool permission guardrail; Portkey MCP and request guardrails; Cloudflare and Kong guardrails on message text | None found | LiteLLM plugin marketplace, Portkey skill sync | Request logs for people: Helicone, LiteLLM, Vercel, Bifrost | None found | LiteLLM falls back to a larger window; OpenRouter context compression | None found | Budgets and logs: LiteLLM, Portkey, Helicone, Cloudflare, Vercel |
| [Token and context optimizers](#optimizers) | Folds on the request path for every machine, and archives the raw output | RTK, Headroom, Caveman, Compresr, LLMLingua; our free Context Mode Engine | Context Mode Engine | RTK Pro command policies | Headroom cross-agent memory | RTK Pro shared skills | Headroom fetches originals back; Context Mode Engine searches its store | Context Mode Engine restores a snapshot after a resume | Compresr summarizes history ahead of compaction | Caveman short replies, Headroom verbosity steering | RTK Pro, `rtk gain` , `headroom savings` |
| [Web and docs retrieval](#web) | A 5-minute cache mark on each WebFetch page; no search or crawl | Firecrawl, Exa, Tavily, Jina Reader and Context7 return shorter page or docs text | None found | None found | None found | Context7 installs a skill | None found | None found | None found | None found | None found |
| [Agent security and MCP gateways](#security) | Cage checks each call before it runs, by what it does | None found | None found | CC Safety Net, dcg, Zenity, Lasso, Knostic, Prisma AIRS, Operant, MintMCP, Obot, Docker MCP Gateway | None found | Snyk Agent Scan and Cisco mcp-scanner scan skills | None found | None found | None found | None found | Audit trails: Zenity, Lasso, MintMCP, Obot |
| [Sandboxes and code execution](#sandboxes) | Thinking in Code runs scripts in a Worker isolate; Cage checks the hosts they call | None found | Cloudflare Code Mode, for agents you build; E2B, Daytona, Modal and Docker Sandboxes run code | Claude Code /sandbox, Codex sandbox, Cursor sandbox, Docker Sandboxes | None found | None found | None found | None found | None found | None found | None found |
| [Memory layers](#memory) | One store for both clients; old facts closed, not deleted | None found | None found | None found | mem0, Zep and Graphiti, supermemory, claude-mem, Basic Memory, Letta Code, Cognee, Headroom | Letta Code | claude-mem, Basic Memory | None found | None found | None found | None found |
| [Session history and resume](#sessions) | One copy on the gateway; resume in either client; search the archive | None found | None found | None found | None found | SpecStory Lore turns history into skills | SpecStory Cloud, Entire, cass, Claude Code History Viewer | Entire, SpecStory, cass | None found | None found | SpecStory Cloud analytics |
| [Observability and team analytics](#observability) | Tokens, cost and saving per request, read from the whole request | Datadog finds waste and suggests a fix | None found | None found | None found | None found | Traces for people: Langfuse, LangSmith, Phoenix, Braintrust | None found | None found | None found | Langfuse, Datadog, Jellyfish, LinearB, DX, Swarmia, Faros AI, ccusage |
| [Native client features](#native) | One set of features across both clients and every machine | Claude Code and Codex compaction, Claude Code tool search, Anthropic context editing, prompt caching | Anthropic programmatic tool calling, Codex `exec` | Claude Code permissions and hooks, Codex approvals and requirements.toml, Cursor run modes | Claude Code auto memory, Codex memories, Anthropic memory tool | Claude Code and Codex skills | None found | `claude --resume` , `codex resume` , `--teleport` from Claude Code on the web | Claude Code and Codex auto-compaction, Anthropic compaction and context editing | Claude Code output styles | Claude Code analytics and OpenTelemetry, Cursor analytics |
| [Skills and rules](#skills) | One catalog on the gateway; at most one skill a turn, named in the reply | None found | None found | Rules files ask; Cursor can make team rules required | CLAUDE.md, AGENTS.md | Agent Skills, skills.sh, anthropics/skills, Cursor, Cline, Roo and Continue rules | None found | None found | None found | A line in CLAUDE.md or AGENTS.md can ask for a style | None found |

## Four questions for each tool

- **Where does it run?** In the client, on your machine, on the request path between the agent and the model, or in a vendor cloud that reads data after the fact.
- **Does it enforce or ask?** A rules file is text the model reads, and the model can still do something else. A check that runs before the tool call cannot be skipped by the model.
- **Which agents?** Claude Code, Codex, Cursor and others.
- **Does it change the request, or only record it?**

## LLM gateways

A gateway sits between the agent and the model provider. Most route requests to many providers, set budgets and keep logs.

| Tool | Runs where | Agents named | License | What it does | What it does not do |
| --- | --- | --- | --- | --- | --- |
| [**Context Mode Gateway**](https://context-mode.com/docs) | Hosted on Cloudflare; one config change per agent | Claude Code and Codex, one policy | Commercial; free to start, then Pro, Team or Enterprise. Your Claude or OpenAI account pays the model, with no fee on tokens | Folds tool output once and archives the raw output, trims tool schemas, keeps each thread under a context limit, checks each tool call with Cage, adds memory, skills and a reply style. [65.5% lower cost](https://context-mode.com/benchmarks/long-session/ab.json) on a long Claude Code session. | Routes to Anthropic and OpenAI, not to other providers. Hosted only. No SSO or team roles yet. |
| [LiteLLM](https://docs.litellm.ai/docs/proxy/guardrails/tool_permission) | Self-hosted proxy: Docker, Kubernetes, AWS or GCP | Claude Code, Codex CLI, Claude Desktop, any client of the OpenAI or Anthropic API format | MIT core; its `enterprise` folder is under a commercial license | Routes to 100+ providers. Virtual keys, budgets and logs. A tool permission guardrail allows or denies tool calls by name and checks argument values; with `on_disallowed_action` , Block halts the request and Rewrite strips the forbidden tools and returns an error message inside the response. An MCP gateway and a Claude Code plugin marketplace. Falls back to a model with a larger window when the context is full. | Folding tool output, memory and session resume are not documented. |
| [Portkey](https://portkey.ai/docs/product/guardrails) , now Prisma AIRS AI Gateway | Hosted; its open-source gateway also runs on Node, Docker, Workers or Kubernetes | Claude Code, Codex | Gateway MIT; platform commercial. Palo Alto Networks bought Portkey in May 2026. | 1,600+ models, budgets, fallbacks, caching, RBAC, an MCP gateway. MCP guardrails check tool-call arguments and results; a Request Parameters Check allows or blocks the tools a request offers. Syncs team skills to SKILL.md files. | Its guardrails on the model request check the last message. Checks on built-in shell tool calls, folding, memory and resume are not documented. |
| [Helicone](https://docs.helicone.ai/integrations/anthropic/claude-code) | Cloud, or self-hosted with Docker or Helm | Claude Code | Apache 2.0. Mintlify bought Helicone. [Its post](https://www.helicone.ai/blog/joining-mintlify) of Mar 3, 2026 says its services stay live in maintenance mode, and security updates, new models, and bug and performance fixes all keep shipping. | Request logs, cost, sessions, caching, rate limits, prompts. | Its sessions group requests for viewing; they do not resume. Its Claude Code setup gets no new features. Tool-call policy is not documented. |
| [OpenRouter](https://openrouter.ai/docs/features/message-transforms) | Cloud | Claude Code (a setup guide) | Commercial | Failover between providers, org budgets, analytics. Its context-compression plugin removes or cuts messages from the middle of the prompt until it fits the window; it is on by default for endpoints with 8k (8,192 tokens) of context or less. | Tool-call policy, an archive of what it drops, and memory are not documented. |
| [Cloudflare AI Gateway](https://developers.cloudflare.com/ai-gateway/) | Cloudflare's network | Claude Code, through Anthropic, Bedrock or Vertex | Commercial, on all plans | Analytics, logs, cache, rate limits, retries and fallbacks, BYOK. Guardrails check prompts and responses for unsafe content, then flag or block. | Checks on tool calls and folding are not documented. |
| [Kong AI Gateway](https://developer.konghq.com/ai-gateway/) | Konnect control plane; data plane self-hosted or on Kubernetes | Claude Code, Codex CLI, Qwen Code | Commercial | Routing, semantic cache, token rate limits, prompt guard, PII redaction, MCP and A2A. A prompt compressor, on the Enterprise tier, uses LLMLingua 2 and drops text. | Memory, resume and tool-call rules that read shell commands are not documented. |
| [Vercel AI Gateway](https://vercel.com/docs/ai-gateway/coding-agents) | Vercel's cloud | About 30 agents through one setup command; Claude Code and Codex have their own endpoints | Commercial; no markup on token prices | Routing, fallbacks, budgets, request logs, BYOK. Works with a Claude subscription. | Folding, tool-call policy, memory and resume across machines are not documented. |
| [Claude apps gateway](https://code.claude.com/docs/en/claude-apps-gateway) | Your own cloud: Kubernetes, Cloud Run, ECS | Claude Code and Claude Desktop | First-party (Anthropic); built into the `claude` binary | SSO by identity-provider group, model allowlists, managed settings for each group (permission rules, hooks, an egress allowlist), spend limits, OpenTelemetry, audit events, failover across Bedrock, Vertex, Foundry and the Anthropic API. | No Codex. Its policy runs in the client through managed settings. Rewriting requests is not documented. |
| [Bifrost](https://docs.getbifrost.ai/cli-agents/overview) | Local or self-hosted. Its Edge version, in alpha, runs on each company machine. | Claude Code, Codex CLI, Gemini CLI, Cursor, Copilot and others | Apache 2.0, plus an enterprise edition | 23+ providers, virtual keys, budgets, semantic cache, an MCP gateway, guardrails, audit logs. | Tool-call policy and folding are not documented. |
| [agentgateway](https://github.com/agentgateway/agentgateway) | Standalone or on Kubernetes | Client guides for Claude Code, Claude Desktop, Codex, Cursor, Copilot and others | Apache 2.0 (Linux Foundation) | A proxy for LLM, MCP and A2A traffic, with RBAC and a CEL policy engine. | Folding, memory and shell-command policy are not documented. |
| [claude-code-router](https://github.com/musistudio/claude-code-router) | Your machine | Claude Code, Codex, Grok CLI, Kimi CLI, OpenCode and others | MIT | Sends each agent to the models and providers you pick, with fallbacks and request logs. | Compression, tool-call policy and team policy are not documented. |
| [Requesty](https://docs.requesty.ai/) | Cloud, with EU routing | Claude Code (a setup guide) | A hosted service; license not documented | Routing, fallbacks, load balancing, caching, spend limits, guardrails, SSO, RBAC and an MCP gateway. | Folding, memory and shell-command policy are not documented. |
| [TrueFoundry AI Gateway](https://www.truefoundry.com/ai-gateway) | Your VPC, on-prem, air-gapped or across clouds | Not documented on the page we read | Commercial | 1,600+ models, policy control, real-time monitoring and failover, SSO, RBAC and audit logs, plus an MCP gateway, an agent gateway and an agent skills registry. | Coding-agent setup, folding and shell-command policy are not documented on the page we read. |
| [Azure API Management AI gateway](https://learn.microsoft.com/en-us/azure/api-management/genai-gateway-capabilities) | Azure | Not documented on the page we read | Commercial | Token limits and quotas per API consumer, semantic caching, content safety checks, remote MCP servers and A2A APIs. | Folding and tool-call policy are not documented. |
| [Agent Router](https://aigateway.envoyproxy.io/) , formerly Envoy AI Gateway | Self-hosted on Envoy, or as a service hosted by Tetrate | Any OpenAI-compatible client | [Apache 2.0](https://github.com/envoyproxy/ai-gateway) (Agentic AI Foundation) | One API for models and one router for MCP tools, with credentials, limits and failover. | Folding, memory and tool-call policy are not documented. |

Left out: [TensorZero](https://github.com/tensorzero/tensorzero), whose repository was archived on 2026-06-12, and Martian and Unify, which now sell other products.

### Where Context Mode differs

- **It changes what the request carries.** Most of these gateways route, cache and log; OpenRouter and Kong can shorten a prompt to fit. Context Mode folds tool output the first time the model sees it and sends it the same way on later turns, so the prompt cache keeps reading it. It trims tool schemas and archives the raw output. See [Context Saving](https://context-mode.com/docs/how-saving-works).
- **Tool calls are read as shell commands.** LiteLLM's tool permission guardrail matches tool names and argument values, and Portkey's MCP guardrails check the arguments of MCP tool calls. Cage reads a shell command through about 130 wrappers and carriers, such as `sudo`, `sh -c`, `xargs` and `ssh`. See [Cage](https://context-mode.com/docs/cage).
- **One set of features for two clients.** Many gateways route both Claude Code and Codex. Context Mode applies folding, Cage, memory, skills, the reply style and the context limit to both.
- **A saving you can check.** On a long session on Opus 5.5, [65.5% lower cost](https://context-mode.com/benchmarks/long-session/ab.json) than Claude Code going direct, from [$7.20](https://context-mode.com/benchmarks/long-session/ab.json) to [$2.49](https://context-mode.com/benchmarks/long-session/ab.json). 155 of 168 shorter runs were cheaper. See [Benchmarks](https://context-mode.com/docs/benchmarks).
- **No fee on tokens.** Your own account pays the model. OpenRouter lists a 5.5% platform fee on Standard and 8% on Business, Requesty bills spend "plus 5%", and Cloudflare adds 5% to credits bought through Unified Billing (pricing pages, read 2026-09-30).

### When to pick another tool

Choose LiteLLM, Portkey, OpenRouter or the Claude apps gateway instead when you need many model providers, your own cloud, or SSO and roles today.

## Token and context optimizers

These tools cut what the agent reads. Most run on each developer's machine.

| Tool | Runs where | Agents named | License | What it does | What it does not do |
| --- | --- | --- | --- | --- | --- |
| [**Context Mode Gateway**](https://context-mode.com/docs/how-saving-works) | The request path, for every machine pointed at the gateway | Claude Code and Codex | Commercial; in every plan | Folds tool output the first time it passes, so the cache keeps reading it, and archives the raw output. The agent reads it back byte for byte, images aside, with secrets and e-mail addresses masked. Paired runs on Opus 5.5 with every feature on: [65.5% lower cost](https://context-mode.com/benchmarks/long-session/ab.json) on a long session. We have not yet split out the share that folding makes. | Run with no network. Trim tool schemas for Codex yet, worth at most about 3k cached tokens a request. |
| [RTK](https://github.com/rtk-ai/rtk) | A binary and an agent hook on each machine | Claude Code, Codex, Cursor, Gemini CLI, Windsurf, Cline and others; its README says it supports 18 AI coding tools | Apache 2.0 | Rewrites shell commands to `rtk` versions that filter, group and cut output, for 100+ commands. `rtk gain` shows estimated savings. | Works on shell output. Its README says token counts are estimates. Saves the full output to a local SQLite store when a command fails or is cut; the agent reads it with `rtk recall` . Output from successful runs is kept only if you turn that on. |
| [RTK Pro](https://pro.rtk-ai.app/) | A SaaS tenant, your own cloud (Helm) or on-prem | "Every agent in your organization"; no list | Commercial; [Dev $6 a month for 1 seat](https://pro.rtk-ai.app/pricing) , Teams from $100 a month for 5 seats, Business from $750 a month for 50 seats, Enterprise custom; a 7-day free trial on Dev and Teams | Secret rewriting, token views by team and project, org-wide policies that block dangerous commands, shared skills, sandboxes. | A full-output archive and measured savings are not documented. |
| [Headroom](https://github.com/headroomlabs-ai/headroom) | Your machine, as a library, a local proxy or an MCP server; a paid option for teams | Claude Code, Codex, Cursor, Copilot, Cline, Aider, Goose and others | Apache 2.0 | Compresses tool output, logs, files and history. Keeps the originals on your machine and fetches them on demand. Leaves the cached prefix unchanged and compresses only new bytes. Also trims output tokens and keeps a memory store shared across agents. Its README reports 21 to 57% fewer tokens on four scenarios built from MCP output, and 90% on repeated JSON and log lines. | Tool-call policy and memory shared across machines are not documented. Its savings are token counts, not paired cost runs at list price. |
| [Caveman](https://github.com/JuliusBrussee/caveman) | Your machine: a skill, a local proxy, middleware | 30+ agents, among them Claude Code, Codex, Cursor and Gemini | Skill and CLI MIT; proxy BSL 1.1 | Makes the agent answer in short text. The proxy shrinks what the agent reads. | Its README says no reviewed API benchmark is published yet. |
| [Compresr Context Gateway](https://github.com/Compresr-ai/Context-Gateway) | A local proxy | Claude Code, Cursor, OpenClaw | Apache 2.0 | Builds the history summary in the background, so compaction is instant. | Measured savings, policy and recovery of dropped text are not documented. |
| [LLMLingua](https://github.com/microsoft/LLMLingua) | A library in your code | None directly; works with LangChain and LlamaIndex | MIT (Microsoft) | Compresses prompts by dropping tokens the model can do without. | A coding-agent setup is not documented. The text it drops is gone. |
| [Augment Context Engine](https://www.augmentcode.com/context-engine) | An MCP server, local or hosted | Claude Code, Codex, Cursor, Gemini CLI, Copilot and others | Commercial | Code search by meaning across repos, history and docs, so the agent reads less. Reports 32% fewer tokens on its own benchmark. | Does not change the rest of the request. Its figure is a benchmark comparison, not a paired cost run. |
| [Claude Context](https://github.com/zilliztech/claude-context) | An MCP server, with Milvus or Zilliz Cloud and an embedding API | Claude Code, Codex CLI and other MCP clients | MIT | Code search by meaning. Its README reports about 40% fewer tokens for equal retrieval quality. | Needs a vector database and an embedding key. Retrieval only. |
| [Serena](https://github.com/oraios/serena) | An MCP server on your machine, using language servers | Any MCP client | Open source; a paid JetBrains plugin | Finds and edits code by symbol instead of reading whole files. | A measured cost figure is not documented. |
| [Repomix](https://github.com/yamadashy/repomix) | A CLI, a website or an MCP server | Any LLM | Open source | Packs a repository into one file for a model, with an optional compressed form and token counts per file. | Works on repository input only; tool output and history are out of scope. |

### Where Context Mode differs

- **On the request path.** Every request through the gateway is folded, from Claude Code or Codex, on any machine that points at it. Nothing keeps running on the machine.
- **Folded once, so the cache holds.** A result is folded the first time it passes. Anthropic's docs say that clearing old tool results later costs [cache writes](https://platform.claude.com/docs/en/build-with-claude/context-editing). Headroom also leaves the cached prefix unchanged.
- **Folded output is archived.** The raw output goes to your archive before it is shortened, and each cut says what was dropped. The agent reads it back byte for byte, images aside, with secrets and e-mail addresses masked, from any machine. Headroom and RTK also keep originals, on the machine that made them.
- **Measured in dollars.** On a long session on Opus 5.5, [65.5% lower cost](https://context-mode.com/benchmarks/long-session/ab.json) than Claude Code going direct, from [$7.20](https://context-mode.com/benchmarks/long-session/ab.json) to [$2.49](https://context-mode.com/benchmarks/long-session/ab.json), 4 of 4 pairs cheaper. We have not yet split out the share that folding makes. Headroom and Augment publish token counts. See [Benchmarks](https://context-mode.com/docs/benchmarks).

### When to pick another tool

Choose RTK or Headroom instead when nothing may leave the machine or you use agents beyond Claude Code and Codex; one person on one machine can start with our free [Context Mode Engine](https://context-mode.com/engine).

### Context Mode Engine among local optimizers

runs where Your machine: an MCP server plus hooks in the client. Tools without hooks use a routing file such as AGENTS.md.

agents 18 AI coding tools, among them Claude Code, Cursor, Copilot, Codex and Gemini CLI.

license Free, source-available under the Elastic License 2.0. No account. 24,230 GitHub stars on 2026-09-30.

what it does Runs large tool output in a sandbox, indexes it in a local SQLite FTS5 store the agent searches, lets the agent write a short script so only its output enters the context, and restores a snapshot after compaction or a resume.

does not Enforce one policy across machines or keep one store across machines. The gateway does those for Claude Code and Codex.

RTK and Headroom are its closest neighbours. RTK rewrites shell commands; Headroom compresses through a local proxy. Context Mode Engine works through MCP tools and hooks, and keeps the full output in a store the agent can search. See [Context Mode Engine](https://context-mode.com/engine).

## Web and docs retrieval

These tools hand the agent shorter text for a web page or a library's docs, so it reads less.

| Tool | Runs where | Agents named | License | What it does | What it does not do |
| --- | --- | --- | --- | --- | --- |
| [**Context Mode Gateway**](https://context-mode.com/docs/how-saving-works#page-cache) | The request path | Claude Code | Commercial | Marks each Claude Code WebFetch page for a 5-minute cache, so the next read of that page within 5 minutes is a cache read. The first read costs 1.25 times a fresh read. | Search or crawl the web. Use these tools with it. |
| [Firecrawl](https://github.com/firecrawl/firecrawl) | A hosted API, or self-hosted | Claude Code, Antigravity, OpenCode and MCP clients | AGPL-3.0; the cloud has more features | Search, scrape and crawl, returned as clean Markdown or structured data. | Changing the rest of the request and tool-call policy are out of scope. |
| [Exa](https://github.com/exa-labs/exa-mcp-server) | A hosted search API and MCP server | Codex, Claude (a plugin) and MCP clients | MCP server MIT; the API is commercial | Web search with highlights, summaries and page crawling. | Changing the rest of the request is out of scope. |
| [Tavily](https://github.com/tavily-ai/tavily-mcp) | A hosted API and MCP server, remote or local | Claude Code, Cline and MCP clients | MCP server MIT; the API needs a key | Search, extract, map and crawl tools. | Changing the rest of the request is out of scope. |
| [Jina Reader](https://github.com/jina-ai/reader) | A hosted API: a URL prefix | Any agent that can fetch a URL | Apache 2.0 | Turns a URL into LLM-friendly text, and a query into search results. | Changing the rest of the request is out of scope. |
| [Context7](https://github.com/upstash/context7) | A hosted service, through a CLI with a skill or an MCP server | Cursor, Claude Code and other coding agents | Client MIT; its API, parser and crawler are private | Up-to-date library docs, fetched when the agent needs them. | Covers library docs only. |

### Where Context Mode differs

- **It does not search or crawl.** Use these tools with it. The gateway gives each Claude Code WebFetch page a 5-minute cache mark. See [the page cache](https://context-mode.com/docs/how-saving-works#page-cache).

### When to pick another tool

Choose Firecrawl, Exa, Tavily, Jina Reader or Context7 when the agent must search, crawl or read library docs; they work next to the gateway.

## Agent security and MCP gateways

These tools decide what a coding agent may run, reach or read. Some hook into the client, some proxy MCP traffic, and some scan before use.

| Tool | Runs where | Agents named | License | What it does | What it does not do |
| --- | --- | --- | --- | --- | --- |
| [**Context Mode Cage**](https://context-mode.com/docs/cage) | The gateway, before Claude Code or Codex runs the call | Claude Code and Codex, one policy | Commercial; in every plan | Reads each call by what it does: 22 capabilities, through 98 wrapper words and 33 carriers. Monitor mode. An API key can only tighten the policy. Log export as CSV or JSON Lines; a hash-chained policy history. | Ask a person first. SSO or group scope. See inside a script file or a compiled program; keep your sandbox on. |
| [CC Safety Net](https://github.com/kenryu42/cc-safety-net) | A hook on each machine | Claude Code, Codex, Cursor, Gemini CLI, GitHub Copilot CLI, OpenCode and others | MIT | Parses what a command does, so a wrapper or reordered flags do not hide it. Blocks SSH keys, `.env` files and the credential files coding CLIs keep. Rulebooks for Terraform, AWS, gcloud and Azure. | Each machine installs the hook; the policy is shared through git. Its audit trail stays on the machine. |
| [dcg](https://github.com/Dicklesworthstone/destructive_command_guard) | A hook on each machine | Claude Code, Codex CLI, Gemini CLI, Copilot CLI, Cursor and others | A custom license based on MIT, with an OpenAI/Anthropic rider | 50+ packs for databases, Kubernetes, Docker, clouds and Terraform. Scans heredocs and inline scripts. An `ask` rule where the hook protocol supports it. | Config lives on each machine. Malformed hook input is allowed with an audit warning unless you opt into fail-closed. |
| [Zenity](https://www.zenity.io/use-cases/agent-type/coding-personal-agents) | Hooks on each machine, and an MCP gateway | Claude Code, Codex, Copilot, Cursor | Commercial | Blocks inline, one central policy, audit through hooks and OpenTelemetry. | Changing requests or cost is not documented. |
| [Lasso](https://www.lasso.security/use-cases/ai-coding-assistants) | Client hooks, rolled out with managed settings; scanning in Lasso's cloud | Claude Code, Cursor, Codex, OpenCode | Commercial; its [Claude Code hook](https://github.com/lasso-security/claude-hooks) is MIT | Scans content for injected instructions, checks tool calls before they run, flags or blocks, keeps an audit trail. | Its open-source hook warns but does not block. Changing requests or cost is not documented. |
| [Knostic Kirin](https://www.knostic.ai/ai-coding-security-solution-kirin) | An IDE extension | Copilot, Cursor, Claude Code, Windsurf | Commercial | Blocks risky parts and unsafe actions, with a central audit. | Codex is not named. |
| [Prisma AIRS](https://www.paloaltonetworks.com/prisma/prisma-ai-runtime-security) (Palo Alto Networks) | Endpoint, network and cloud; an inline AI gateway from Portkey | Cursor, Claude Code, Codex, Antigravity | Commercial | One policy across agents, blocks with an exception request, session timelines. The gateway covers LLM, MCP and A2A traffic. | Context savings and an archive are not documented. |
| [Prompt Security](https://www.sentinelone.com/press/sentinelone-to-acquire-prompt-security/) (SentinelOne) | A commercial platform with an MCP gateway | Names ChatGPT, Gemini and Claude | Commercial | Visibility, blocks prompt injection and data leaks. Its MCP gateway covers 13,000+ known MCP servers. | Checks on coding-agent tool calls are not documented in the release. |
| [Lakera Guard](https://docs.lakera.ai/docs/quickstart) (Check Point) | A screening API | Any app | Commercial | Screens prompts and outputs for injection, personal data and bad links. | Checks on coding-agent tool calls are not documented. |
| [Operant AI](https://operant.ai) | An MCP gateway and endpoint agent | Claude Code, Cursor | Commercial | Allows, blocks or redacts in real time. | Codex is not named. |
| [MintMCP](https://mintmcp.com) | A hosted MCP gateway and agent monitor | Claude Code, Cursor, ChatGPT | Commercial | RBAC, audit, scanning for personal data and secrets. | Covers MCP and tool calls only. |
| [Obot](https://github.com/obot-platform/obot) | An MCP gateway, self-hosted or cloud | Claude, Cursor, VS Code | MIT | RBAC, audit of tool calls and LLM requests, finds MCP servers nobody approved. | Shell commands are out of scope. |
| [Docker MCP Gateway](https://docs.docker.com/ai/mcp-catalog-and-toolkit/mcp-gateway/) | Local containers | VS Code, Cursor, Claude Desktop, Claude Code | MIT | Runs each MCP server in its own container and manages its credentials and OAuth. | Covers MCP traffic only; shell commands and model requests are out of scope. |
| [Snyk Agent Scan](https://github.com/snyk/agent-scan) | A local CLI; needs a Snyk token | Claude Code, Codex, Cursor, VS Code, Copilot, Gemini CLI and more | Apache 2.0 | Scores MCP servers and skills for prompt injection, destructive tools and malware. | Scans only; it does not block at run time. |
| [Cisco mcp-scanner](https://github.com/cisco-ai-defense/mcp-scanner) | A local CLI, an API or a gateway | Reads the configs of Cursor, Claude, VS Code and Windsurf | Apache 2.0 | Scans MCP servers with YARA rules, an LLM and an API. | Scans only. |
| [Invariant Gateway](https://github.com/invariantlabs-ai/invariant-gateway) | A proxy, hosted or self-hosted | Frameworks such as Swarm, Autogen, OpenHands and SWE-agent | Apache 2.0 | Traces, and guardrails that block or log. | Claude Code and Codex are not named. Last change on GitHub: 2025-11-06. |

### Where Context Mode differs

- **The check is on the request path.** Once you turn Cage on, it decides before Claude Code or Codex runs a call. A blocked command never runs; the terminal runs only an echo that carries the refusal. On each machine it depends on one config change, which the CLI writes. A machine without it is outside Cage: its calls never reach the gateway.
- **One policy for both clients, with no install.** The same rules cover Claude Code and Codex on every machine pointed at the gateway. CC Safety Net and dcg read commands well too, with a hook on each machine and a log on each machine.
- **The same hop saves and records.** The hop that checks the call also folds the output, archives it and measures the saving.
- **An audit trail you can check.** Every block goes to a decision log you can export. Each policy change records who, how and why, and an edited or deleted entry breaks the hash chain.

### When to pick another tool

Choose CC Safety Net or dcg instead when you want a free check on one machine with no network, and Prisma AIRS, Zenity or Lasso when you need SSO, group policies or models that detect injection.

## Sandboxes and code execution

A sandbox limits what a program can reach once it runs: files, network, the rest of the machine. Code execution for agents runs code the model writes, so only the result enters the context.

| Tool | Runs where | Agents named | License | What it does | What it does not do |
| --- | --- | --- | --- | --- | --- |
| [**Context Mode Thinking in Code and Cage**](https://context-mode.com/thinking-in-code) | Cloudflare Worker isolates for the agent's scripts; the gateway for Cage | Claude Code and Codex, with no code change | Commercial | The agent writes a short script; it runs in a Worker isolate, and what it prints reaches the model. With Cage on, every host the script calls is checked. | Isolate your machine. Run long jobs or other languages: scripts are short JavaScript, and local files go through your agent's shell. |
| [Claude Code /sandbox](https://code.claude.com/docs/en/sandboxing) | Your machine: Seatbelt on macOS, bubblewrap and seccomp on Linux | Claude Code; its [runtime](https://github.com/anthropics/sandbox-runtime) can wrap any process | Runtime Apache 2.0 | File and network limits through a local proxy with a domain allowlist. Admins can require it. | Anthropic's page says it is "not a complete isolation boundary". Without its dependencies it runs commands outside the sandbox unless an admin turns that off. |
| [Codex sandbox](https://developers.openai.com/codex/concepts/sandboxing.md) | Your machine: Seatbelt, bubblewrap and seccomp, a Windows sandbox | Codex | Apache 2.0 | Network off by default. Admins lock sandbox modes, approval policies and domain rules in requirements.toml. | Its settings do not limit web search, browser tools, MCP connections or model API calls. |
| [Cursor sandbox and run modes](https://cursor.com/docs/agent/security/run-modes) | Your machine: Seatbelt on macOS, Landlock and seccomp on Linux | Cursor | Commercial | Auto-review, allowlist and run-everything modes, network rules in `sandbox.json` , admin overrides from the dashboard. | Cursor says its auto-review classifier "can make mistakes". Cursor only. |
| [Docker Sandboxes](https://docs.docker.com/ai/sandboxes/) | A microVM, local or in the cloud | Claude Code, Codex, Copilot, Cursor, Gemini, Kiro, OpenCode and others | Commercial; org policy is on a paid plan | Isolates the agent in a microVM. A host proxy applies network policy and injects credentials. | Docker's page says "the agent has full control inside the VM". Single tool calls are not read. |
| [Claude Code dev container](https://code.claude.com/docs/en/devcontainer) | Docker, local or Codespaces | Claude Code | A reference setup in an open-source repo | An isolated environment with a reference firewall script. | With permission prompts skipped, Anthropic says it does not stop a malicious project from sending out what the container can read. |
| [E2B](https://github.com/e2b-dev/E2B) | Cloud Linux VMs, or self-hosted with Terraform | Any agent, through its SDK | Apache 2.0 | Starts a VM on demand to run agent code. | Policy on the tool calls of an agent on a laptop is not documented. |
| [Daytona](https://www.daytona.io/docs/en/) | Cloud | Any agent, through its SDK | AGPL-3.0 (LICENSE at v0.190.0); core development private since June 2026 | Sandboxes with their own kernel, file system and network stack. | Policy on the tool calls of an agent on a laptop is not documented. |
| [Modal Sandboxes](https://modal.com/docs/guide/sandboxes) | Modal's cloud, on gVisor | Any agent, through its SDK | Commercial | Runs code with the network blocked or limited to CIDR ranges or domains. | A sandbox lives 5 minutes by default and 24 hours at most. No tool-call policy. |
| [Cloudflare Code Mode](https://blog.cloudflare.com/code-mode/) | Cloudflare Workers: the model's code runs in a sandbox cut off from the Internet, reaching only the MCP servers it is given | Agents you build with the Cloudflare Agents SDK | Part of the Agents SDK | Turns MCP tools into a TypeScript API. The model writes code against it, and only what the code logs with `console.log` goes back to the agent. Published 2025-09-26. | For agents you build. Setup for Claude Code or Codex is not documented in the post. |

### Where Context Mode differs

- **Thinking in Code for the agents you already use.** The agent writes a short script, the gateway runs it in a Cloudflare Worker isolate, and only what it prints reaches the model. Cloudflare Code Mode uses the same pattern for agents you build with the Agents SDK. The gateway gives the pattern to Claude Code with no code change. Codex has its own `exec`, which runs a JavaScript program on your machine; through the gateway, the script runs off the machine in a Worker isolate. When Cage is on, every host the script calls is checked against it. See [Thinking in Code](https://context-mode.com/thinking-in-code).
- **Cage decides before the call runs.** A sandbox limits a program once it runs. Cage stops the call from being run, and the block is in the decision log.
- **One policy for two clients.** The client sandboxes are set up per client and per machine.

### When to pick another tool

Keep your Claude Code, Codex or Cursor sandbox on in every case, and choose E2B, Daytona, Modal or Docker Sandboxes when a job needs a full Linux machine: a sandbox limits a program once it runs, and Cage reads the call before it runs.

## Memory layers

These tools carry what an agent learned into later sessions. Some write files on your machine, some are APIs you call from your own code.

| Tool | Runs where | Agents named | License | What it does | What it does not do |
| --- | --- | --- | --- | --- | --- |
| [**Context Mode Memory**](https://context-mode.com/docs/memory) | The gateway: one store per account, nothing on each machine | Claude Code and Codex share one store | Commercial; in every plan | Picks facts from turns and tools, and adds them when a conversation starts or after compaction. A changed fact closes the old row. In a controlled test on 24 facts: 0 of 24 blocks held an old value, against 23 of 24 for an append-only store; 38 of 38 "what was true then" answers. | A single "as of" query yet. Team memory apart from personal memory. |
| [Claude Code memory](https://code.claude.com/docs/en/memory) | Local files: CLAUDE.md you write, and auto memory notes under `~/.claude/projects` with a MEMORY.md index loaded each session | Claude Code | Part of the client | Instructions you write, plus notes Claude writes as it works. | Stays on one machine. Codex does not read the auto memory notes. |
| [Codex memories](https://developers.openai.com/codex/memories) | Local files under `~/.codex/memories/` | Codex | Part of the client | Summaries and entries from earlier chats. Off by default. | OpenAI's page says to keep required team guidance in AGENTS.md and treat memories as a recall layer. Stays on one machine. |
| [Anthropic memory tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/memory-tool) | Your application: Claude asks for file operations under `/memories` , and your code runs them against storage you pick | Apps built on the Claude API | Part of the API | A building block for your own agent's memory. | You write and host the storage. Not a feature of Claude Code or Codex. |
| [mem0](https://github.com/mem0ai/mem0) and [OpenMemory](https://mem0.ai/openmemory) | A library, a self-hosted server, or mem0's cloud | OpenMemory: Cursor, VS Code, Claude and MCP clients | Apache 2.0; cloud is commercial | Captures coding preferences and patterns as you work and adds the ones that match the current project. | Tool-call policy and output folding are not documented. |
| [Graphiti](https://github.com/getzep/graphiti) and Zep | Graphiti in your own graph database; Zep as a managed service | Any MCP client, through its MCP server | Graphiti Apache 2.0; Zep commercial | A graph of facts over time. A changed fact is marked invalid, not deleted. | Needs an LLM and a graph database. A coding-agent plugin is not documented. |
| [supermemory](https://supermemory.ai/docs/integrations/claude-code) | Cloud, or self-hosted | Claude Code, Codex, Cursor, OpenCode and others, by plugin or MCP | [MIT](https://github.com/supermemoryai/supermemory) | Hooks save conversations and important tool use. Extracts facts, handles changes and contradictions, and keeps team memory apart from personal memory. | Tool-call policy and output folding are not documented. |
| [claude-mem](https://github.com/thedotmack/claude-mem) | Hooks and a local worker with SQLite and Chroma; a hosted tier | Built for Claude Code; installers for OpenCode, Antigravity and others | Apache 2.0 | Saves observations from each session, compresses them, and searches them through MCP tools. | Tool-call policy and output folding are not documented. |
| [Basic Memory](https://github.com/basicmachines-co/basic-memory) | Markdown files and a local SQLite index; a hosted version | Claude, Codex, Cursor, ChatGPT and MCP clients | AGPL 3.0 | Notes you and the agent both read and write, with search and links between them. | Tool-call policy and output folding are not documented. |
| [Letta Code](https://docs.letta.com/letta-code/memory) | Its own agent: a CLI, a desktop app, the browser | Letta Code agents | Apache 2.0 | A git-backed memory file system. Agents rewrite their own memory and skills over time. | A separate agent, not a layer for Claude Code or Codex. |
| [Cognee](https://github.com/topoteretes/cognee) | Local, self-hosted, or Cognee Cloud | Claude Code and Codex, through plugins; MCP clients | Apache 2.0; the cloud is commercial | Turns documents, code and conversations into a knowledge graph the agent searches. Runs locally on small models with no API key. | Tool-call policy and output folding are not documented. |
| [Headroom](https://github.com/headroomlabs-ai/headroom) memory | Your machine, as part of Headroom | Claude, Codex, Gemini and Grok | Apache 2.0 | One shared memory store across agents, with automatic dedup. `headroom learn` writes corrections into CLAUDE.md or AGENTS.md. | A store shared across machines is not documented. |

### Where Context Mode differs

- **Added on the request path.** No plugin, hook or MCP server on each machine. The gateway adds memory when a conversation starts or after compaction.
- **One store per account, on every machine.** Claude Code and Codex share it. Headroom also shares one store across agents, on the machine that runs it.
- **Old facts are closed, not deleted.** A replaced fact stays as history. Graphiti does this too.
- **Measured against the alternatives.** In a controlled test on 24 facts with the gateway's own code, closing old rows kept old values out of 24 of 24 memory blocks, where an append-only store let them into 23; an overwriting store could not answer 38 "what was true then" questions. [Memory](https://context-mode.com/docs/memory#evaluation) gives the live check, cost included.

### When to pick another tool

Choose Claude Code memory or Codex memories instead when you use one client on one machine, and Graphiti, mem0 or supermemory when you need self-hosting, team memory or an API for your own apps.

## Observability and team analytics

These tools record what coding agents do. Tracing tools record it through hooks or OpenTelemetry. Engineering analytics join it to git and tickets.

| Tool | Runs where | Agents named | License | What it does | What it does not do |
| --- | --- | --- | --- | --- | --- |
| [**Context Mode Gateway**](https://context-mode.com/docs/how-saving-works) | The request path; one record for both clients | Claude Code and Codex | Commercial; Analytics is in every plan | Tokens, cost and saving for each request, read from the whole request: system prompt, tool schemas, CLAUDE.md, skills and tool output. Every Cage block. | Evals, datasets or prompt management. Read git or tickets. On Codex, part of the saving is not counted yet. |
| [Langfuse](https://langfuse.com/coding-agents) | Cloud or self-hosted. Hooks or OpenTelemetry on each machine, or behind an LLM gateway. | Claude Code, Codex, Copilot, Cursor, Kiro, OpenCode, Augment Code, VS Code | MIT, except its `ee` folders. Now part of ClickHouse. | Traces, evals, prompt management, cost per person or repo. | Hook tracing does not change requests. Its guide says to treat it "as telemetry, not enforcement", and that its hooks cannot read what CLAUDE.md, skills or auto-loaded context added to the prompt. |
| [LangSmith](https://docs.langchain.com/langsmith/trace-claude-code) | Cloud, hybrid or self-hosted. A Claude Code plugin on each machine. | Claude Code | Commercial | Traces of messages, tool calls, compaction and subagents. Redacts secrets on your machine before upload. Evals. | System prompts are not traced. Does not change requests. |
| [Arize Phoenix](https://arize.com/docs/phoenix/integrations/coding-agents/claude-code) | Self-hosted or cloud. A Claude Code plugin with hooks; OpenTelemetry. | Claude Code, Claude Agent SDK | Elastic License 2.0 | Traces, evals, datasets, experiments, a prompt playground. | Does not proxy requests. |
| [Braintrust](https://www.braintrust.dev/docs/integrations/sdk-integrations/claude-code) | Cloud, or hybrid with your own data plane. A gateway in public preview. | Claude Code | Commercial; its proxy code is open source | Logs, evals, playgrounds. The gateway caches, logs and fails over. | Changing tool output or checking tool calls is not documented. |
| [Datadog Agent Console](https://docs.datadoghq.com/ai_agents_console/) | Datadog's cloud, from Claude Code's OpenTelemetry or the Anthropic usage integration | Claude Code, Cursor, GitHub Copilot | Commercial; in Preview | Spend, sessions, time to merge. Finds problems such as retry loops and reading the same file again, each with a monthly cost and a suggested fix. | Blocking or changing agent requests is not documented. |
| [Jellyfish AI Impact](https://jellyfish.co/platform/ai-impact/) | SaaS | Copilot, Cursor, Claude Code, Amazon Q, Gemini Code Assist, Windsurf, CodeRabbit, Devin, Jules | Commercial | Adoption, spend and delivery for each tool and team, from git and planning tools. | Request content is not among its documented sources. |
| [LinearB](https://linearb.helpdocs.io/article/p4lfzfnu0w-ai-tools-integrations) | SaaS | Copilot, Cursor, Claude Code, Amazon Q Developer, Amazon Kiro | Commercial | Adoption, acceptance and tokens, next to cycle time and change failure rate. | Request content is not among its documented sources. |
| [DX](https://getdx.com/ai-measurement/) | SaaS | Claude Code, Cursor, GitHub Copilot | Commercial | A framework for use, impact and cost. Share of AI code by commit, PR and repo. Unused licenses. | Request content is not among its documented sources. |
| [Swarmia](https://www.swarmia.com/product/ai-impact/) | SaaS | Claude Code, GitHub Copilot, Cursor, Codex, CodeRabbit | Commercial | Cost for each PR and initiative, cycle time, review agents. | Self-hosting is not documented. |
| [Faros AI](https://www.faros.ai/) | SaaS | Not documented on its public pages | Commercial | Links AI coding spend to shipped work, and covers model routing and governance. | Its docs need a login; details are not documented publicly. |
| [ccusage](https://github.com/ccusage/ccusage) | A CLI on your machine, reading local session files | Claude Code, Codex, OpenCode, Amp, Droid and others | MIT | Daily, monthly and per-session token use and cost from local data. | Reads one machine. Does not change requests. |

### Where Context Mode differs

- **It sees the whole request.** The gateway reads what the model reads: system prompt, tool schemas, CLAUDE.md, skills and tool output. Langfuse's guide says its hooks cannot read what CLAUDE.md, skills or auto-loaded context added.
- **It changes the request, then records the change.** Tracers and dashboards record. Datadog finds waste and suggests a fix. Context Mode removes some of that waste on the request path and shows tokens, cost and saving for each request.
- **One record for two clients.** Claude Code and Codex requests go to the same account.

### When to pick another tool

Choose Langfuse, LangSmith, Phoenix or Braintrust instead when you need evals and prompt management, and Jellyfish, LinearB, DX, Swarmia or Faros AI when you want AI use joined to git and tickets.

### Context Mode Insight among team analytics

runs where Our cloud. A developer opts in with an Insight key, and Context Mode Engine forwards each session event: tool name, file path, command type, exit code, and the first 200 characters of the event's text, which includes the start of each prompt. Secrets, e-mail addresses and the home folder name are masked first. File contents stay on the machine.

agents The coding tools Context Mode Engine supports.

license Commercial.

what it does Runs 222 patterns over those events, such as blockers, retry waste and error spikes, and answers through 13 MCP tools, scoped by role: an owner sees the organisation, a manager their teams, a member their own work.

does not Read file contents or change requests. It sees the first 200 characters of each prompt. Link to tickets is not documented.

Jellyfish, LinearB, DX, Swarmia and Faros AI join AI use to git and planning tools. Insight starts from what happens inside agent sessions. See [Context Mode Insight](https://context-mode.com/insight).

## Native client features

Claude Code, Codex and Cursor ship their own tools for context, safety, memory and resume. They are free, first-party and improve with each release.

| Tool | Runs where | Agents | License | What it does | What it does not do |
| --- | --- | --- | --- | --- | --- |
| [**Context Mode Gateway**](https://context-mode.com/docs) | The gateway, outside the client | Claude Code and Codex, checked on 28 gateway mechanisms | Commercial; every plan has every feature | One archive, memory store, skill catalog, reply style, context limit and Cage policy for both clients, on every machine. | Ask a person before a call runs. Work without a third party on the request path. |
| [Claude Code](https://code.claude.com/docs/en/costs) | The client, on your machine | Claude Code | Commercial | Prompt caching, auto-compaction, [MCP tool search](https://code.claude.com/docs/en/mcp) , a cap on MCP output, [permissions and hooks](https://code.claude.com/docs/en/permissions) , managed settings, `--resume` and `--continue` , skills, [analytics](https://code.claude.com/docs/en/analytics) . | The local session file keeps the full transcript; the client gives the model no search over it. Local sessions and memory stay on one machine; sessions run in Claude Code on the web can move to the terminal with [`--teleport`](https://code.claude.com/docs/en/claude-code-on-the-web) . Claude Code only. |
| [Codex](https://learn.chatgpt.com/docs/config-file/config-reference) | The client, on your machine | Codex | Apache 2.0 | [Compaction and `resume`](https://developers.openai.com/codex/cli/slash-commands.md) , automatic compaction, a limit on tool output tokens, memories, sandbox and approvals, admin policy in requirements.toml, skills, the [unified `exec` tool](https://learn.chatgpt.com/docs/config-file/config-reference) that runs commands, and a code mode that is off by default. | The local session file keeps the history; searching it from the model is not documented. Codex only. |
| [Cursor](https://cursor.com/docs/agent/security/run-modes) | The editor, on your machine | Cursor | Commercial | Run modes and a sandbox, [rules](https://cursor.com/docs/context/rules) with required team rules, [team analytics](https://cursor.com/docs/account/teams/analytics) and an Admin API. | Cursor only. |
| [Anthropic context management](https://platform.claude.com/docs/en/build-with-claude/context-editing) | The Claude API | Apps built on the API | Part of the API | Server-side compaction, clearing old tool results and thinking, a memory tool. | For people who build their own agent. Anthropic's docs say clearing tool results costs cache writes. |
| [Anthropic tool search and programmatic tool calling](https://platform.claude.com/docs/en/agents-and-tools/tool-use/programmatic-tool-calling) | The Claude API | Apps built on the API; Claude Code uses tool search | Part of the API | Tools load only when needed. Code runs in a container, so intermediate results stay out of the context. | The app builder has to wire them in. |
| [Prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching) (Anthropic, [OpenAI](https://developers.openai.com/api/docs/guides/prompt-caching) ) | The provider | All clients | Part of the API | Prices repeated input far lower than new input. | Does not shrink what a request carries. |
| [Claude Code output styles](https://code.claude.com/docs/en/output-styles) | The client: Markdown files for the user, the project or managed settings | Claude Code | Part of the client | Changes how Claude Code responds. Built-in styles: Default, Proactive, Concise, Explanatory and Learning, or your own. | Claude Code only. |

### Where Context Mode differs

- **Across clients and machines.** One archive, one memory store, one skill catalog, one reply style, one context limit and one Cage policy for Claude Code and Codex, on every machine.
- **Folded output is archived.** Tool output the gateway folds, and old turns it moves out under the [context limit](https://context-mode.com/docs/context-limit), are archived and can be read back byte for byte, images aside, except secrets and e-mail addresses, which are masked before storage. The gateway keeps each thread under the client's own compaction point.
- **Cage runs outside the client.** On each machine it depends on one config change, which the CLI writes. A machine without it is outside Cage.
- **Thinking in Code off the machine.** Anthropic's programmatic tool calling needs an app you build. The gateway gives the pattern to Claude Code. Codex's own [`exec` tool](https://learn.chatgpt.com/docs/config-file/config-reference) runs on your machine; for Codex, the gateway runs the script off the machine in a Worker isolate.
- **One reply style for both clients.** Claude Code output styles cover Claude Code. The [reply style](https://context-mode.com/docs/reply) is set once and reaches Claude Code and Codex.

### When to pick another tool

Choose the built-in features alone when you use one client, want no third party on the request path, or need an ask step before a call runs.

## Session history and resume

These tools keep what an agent did in a session, so you can search it or pick it up again.

| Tool | Runs where | Agents named | License | What it does | What it does not do |
| --- | --- | --- | --- | --- | --- |
| [**Context Mode Gateway**](https://context-mode.com/docs/session-resume) | The gateway keeps one copy of each conversation | Claude Code and Codex | Commercial | One command rebuilds a conversation as a native Claude Code or Codex session on any machine. [Search](https://context-mode.com/docs/search) finds prompts, commands and archived tool output. | Bring back every tool output: resume returns message text, tool names and a summary. Share sessions with a team. |
| [SpecStory](https://github.com/specstoryai/getspecstory) | Your machine, saving to `.specstory/history/` ; SpecStory Cloud with a login | Claude Code, Codex CLI, Cursor, Copilot, Gemini CLI, Antigravity, OpenCode and others | CLI open source (Apache 2.0 repository); the cloud is commercial | Saves every session locally. With SpecStory Cloud it searches and shares sessions, and resumes them across agents, computers and team members. The cloud also has analytics. | Changing requests and tool-call policy are not documented. |
| [Entire](https://github.com/entireio/cli) | A CLI and git hooks; checkpoints are git refs in your repository | Claude Code, Codex, Cursor, Antigravity, Pi and more | MIT | Saves each agent session with the commit it made. `entire session resume` picks up a stopped session by branch, and checkpoints can be searched. | Changing requests and tool-call policy are not documented. |
| [cass](https://github.com/Dicklesworthstone/coding_agent_session_search) | A TUI and CLI on your machine; syncs other machines over SSH | Codex, Claude Code, Gemini CLI, Cursor, OpenCode, Aider and many more | MIT with an OpenAI and Anthropic rider | One index over local session history from many agents, with resume commands. | Changing requests is not documented. |
| [Claude Code History Viewer](https://github.com/jhlee0409/claude-code-history-viewer) | A desktop app or a headless server, offline | Claude Code, Codex CLI, Gemini CLI, Cursor, Cline, OpenCode and others | MIT | Browses, searches and analyzes local conversation history. | Reads history; resume and request changes are not documented. |
| [Claude Code on the web](https://code.claude.com/docs/en/claude-code-on-the-web) | Anthropic's cloud | Claude Code | Commercial; on Pro, Max and Team plans, and some Enterprise seats | Runs sessions in the cloud and moves them to and from the terminal with `--cloud` and `--teleport` . | Claude Code only. |

### Where Context Mode differs

- **One copy, either client.** The gateway keeps one copy of each conversation that goes through it. One command rebuilds it as a native Claude Code or Codex session on any machine. See [Session resume](https://context-mode.com/docs/session-resume).
- **Search covers the archive.** [Search](https://context-mode.com/docs/search) finds prompts, commands and archived tool output, and the agent reads that output back byte for byte, images aside, except secrets and e-mail addresses, which are masked before storage.
- **Nothing to sync.** The copy is made on the request path, so no local folder, git ref or SSH sync is needed.

### When to pick another tool

Choose built-in resume, SpecStory, Entire or cass instead when you need every tool output of a session back, sessions tied to commits, or agents beyond Claude Code and Codex.

## Skills and rules

Two different things get called rules. A rules file is text the model reads. A permission is a check that runs before the tool call. The model can ignore text. It cannot skip a check.

| Tool | Where it lives | Agents | License | Enforced or advisory | Team-wide |
| --- | --- | --- | --- | --- | --- |
| [**Context Mode Skills**](https://context-mode.com/docs/skills) | The gateway; imported from any GitHub repository with SKILL.md files | Claude Code and Codex | Commercial; in every plan | The gateway picks at most one skill a turn, after a second check, and names it in the reply. Policy is enforced by Cage, not by a rules file. | One catalog per account, on every machine. No team roles yet. |
| [Claude Code skills](https://code.claude.com/docs/en/skills) and [plugins](https://code.claude.com/docs/en/plugins/org) | SKILL.md on the local disk: personal, project, plugin, or the managed settings folder | Claude Code | Part of the client | Guidance. The model loads a skill when its description fits. | Yes. Managed settings deploy skills and require or restrict plugin marketplaces. |
| [Codex AGENTS.md](https://learn.chatgpt.com/docs/agent-configuration/agents-md) and [skills](https://learn.chatgpt.com/docs/build-skills) | AGENTS.md in `~/.codex` and each folder from the git root down; skills in `.agents/skills` and `/etc/codex/skills` | Codex | Part of the client | AGENTS.md files are joined into the prompt, 32 KiB by default. | An admin skills folder, and plugins for rollout across an org. |
| [Cursor rules](https://cursor.com/docs/context/rules) | `.cursor/rules` in git, user rules, team rules in the dashboard, AGENTS.md | Cursor | Commercial | Guidance the model reads. | Team rules can be required for every member. |
| [Cline rules](https://docs.cline.bot/features/cline-rules) | `.clinerules` and a global folder; also reads `.cursorrules` , `.windsurfrules` and AGENTS.md | Cline in VS Code, JetBrains, a CLI and a desktop app | Apache 2.0 | Guidance. | Not documented. |
| [Roo Code rules](https://roocodeinc.github.io/Roo-Code/features/custom-instructions) | `.roo/rules` , rules for each mode, AGENTS.md | Roo Code | Apache 2.0; repository archived on 2026-05-15 | Guidance. | Not documented. |
| [Continue rules](https://docs.continue.dev/customize/deep-dives/rules) | `.continue/rules` | Continue | Apache 2.0; repository no longer actively maintained | Guidance, added to the system message. | Not documented. |
| [Agent Skills](https://agentskills.io) | An open format for skill folders, first developed by Anthropic | Claude Code, Codex, Gemini CLI, OpenCode and others | Open standard | Not applicable. | Not applicable. |
| [skills.sh](https://skills.sh) | A directory by Vercel; installs with `npx skills add` | Claude Code, Cursor, Copilot, Cline, Gemini and others | Not documented | Not applicable. | Not applicable. |
| [anthropics/skills](https://github.com/anthropics/skills) | A GitHub repository that is also a Claude Code plugin marketplace | Claude Code | Apache 2.0; its document skills are source-available | Not applicable. | Not applicable. |

### Where Context Mode differs

- **Skills live on the gateway, not on disk.** One catalog reaches Claude Code and Codex on every machine that goes through the gateway. You import from any GitHub repository that holds SKILL.md files. See [Skills](https://context-mode.com/docs/skills).
- **The gateway picks the skill.** It compares your message with each skill's name and description, runs a second check, loads at most one skill a turn, and names it in the reply. In Claude Code and Codex, the model decides from the descriptions.
- **Policy is a check, not a rules file.** Cage decides on the gateway before Claude Code or Codex runs a call. A line in AGENTS.md or CLAUDE.md is text the model may not follow.
- **The cost is stated.** The skill list was about 760 tokens in our benchmark accounts. [Skills](https://context-mode.com/docs/skills) says what we have not measured.

### When to pick another tool

Choose Claude Code managed settings, plugin marketplaces or the Codex admin skills folder instead when you must roll skills out to a whole org today.

## Limits today

- **Two clients enforced.** Claude Code and Codex. Other agents that go through the gateway are recorded, not enforced.
- **Two providers.** Anthropic and OpenAI.
- **Hosted on Cloudflare, one account.** No on-premises build; on Enterprise, the gateway can run in your own Cloudflare account. No SSO, SCIM or team roles yet; one Cage policy per account.
- **Every prompt passes the gateway.** Secrets and e-mail addresses are masked before the archive, and the Cage log holds no command text. See [what we store](https://context-mode.com/docs/faq#prompts).
- **Data retention.** Stored prompts have no automatic expiry yet, and the console has no delete or export button. Ask us and we delete your account's data; you can export your Cage policy. See [delete or export](https://context-mode.com/docs/faq#delete).
- **Cage reads the call, not the program.** A script file, a compiled program or a payload decoded at run time gets past it. Keep your agent's sandbox on.
- **Codex gaps.** Tool-schema trimming does not run for Codex yet, and part of the saving on Codex is not counted yet. Our paired cost runs are on Claude Code.
- **Tool search needs one setting.** Behind a gateway, Claude Code turns tool search off unless `ENABLE_TOOL_SEARCH=true` is set ([Claude Code MCP docs](https://code.claude.com/docs/en/mcp)). The Context Mode CLI sets it.

## FAQ

### Does Context Mode route to other model providers?

No. Anthropic and OpenAI only.

### Can I self-host it?

Not on your own servers. On Enterprise, the gateway can run in your own Cloudflare account. See [Team and Enterprise](https://context-mode.com/docs/plans-and-limits#team).

### Does it have SSO or team roles?

Not yet.

### Is it an MCP gateway?

No. It is a gateway for the model API. Cage rules also cover calls to tool servers, so one policy covers shell commands, sites, files and MCP tools.

## Sources

Read on 2026-09-29, except the lists marked 2026-09-30, which we read or read again on that day.

- Pricing pages, read 2026-09-30: [Claude](https://claude.com/pricing), [Cursor](https://cursor.com/pricing), [Portkey](https://portkey.ai/pricing), [LiteLLM Enterprise](https://www.litellm.ai/enterprise), [OpenRouter](https://openrouter.ai/pricing), [Requesty](https://www.requesty.ai/pricing), [Cloudflare AI Gateway](https://developers.cloudflare.com/ai-gateway/reference/pricing/), [Helicone](https://www.helicone.ai/pricing), [Kong](https://konghq.com/pricing), [TrueFoundry](https://www.truefoundry.com/pricing), [mem0](https://mem0.ai/pricing), [supermemory](https://supermemory.ai/pricing), [Zep](https://www.getzep.com/pricing), [Letta](https://www.letta.com/pricing), [Langfuse](https://langfuse.com/pricing), [LangSmith](https://www.langchain.com/pricing), [Braintrust](https://www.braintrust.dev/pricing), [Swarmia](https://www.swarmia.com/pricing/), [LinearB](https://linearb.io/pricing), [Snyk](https://snyk.io/plans/), [Docker](https://www.docker.com/pricing/). The ChatGPT, Headroom, RTK Pro, Lasso, Zenity, Prompt Security and Lakera pricing pages were not public or could not be fetched.
- Read again on 2026-09-30: [RTK README](https://github.com/rtk-ai/rtk) (18 AI coding tools, `rtk recall`), [RTK Pro pricing](https://pro.rtk-ai.app/pricing) (Dev $6 a month; 7-day trial), [Helicone joins Mintlify](https://www.helicone.ai/blog/joining-mintlify), [Codex config reference](https://learn.chatgpt.com/docs/config-file/config-reference) (unified `exec` tool, code mode)
- Command guards, read 2026-09-30: [CC Safety Net](https://github.com/kenryu42/cc-safety-net), [dcg](https://github.com/Dicklesworthstone/destructive_command_guard)
- Read on 2026-09-30: [LiteLLM tool permission](https://docs.litellm.ai/docs/proxy/guardrails/tool_permission), [Portkey MCP guardrails](https://portkey.ai/docs/product/mcp-gateway/guardrails), [Portkey guardrail checks](https://portkey.ai/docs/product/guardrails/list-of-guardrail-checks), [agentgateway clients](https://agentgateway.dev/docs/standalone/latest/integrations/llm/clients/claude-code/), [OpenRouter context compression](https://openrouter.ai/docs/features/message-transforms), [Requesty](https://docs.requesty.ai/), [TrueFoundry AI Gateway](https://www.truefoundry.com/ai-gateway), [Azure API Management AI gateway](https://learn.microsoft.com/en-us/azure/api-management/genai-gateway-capabilities), [Agent Router](https://aigateway.envoyproxy.io/), [Agent Router on GitHub](https://github.com/envoyproxy/ai-gateway), [Headroom](https://github.com/headroomlabs-ai/headroom), [Firecrawl](https://github.com/firecrawl/firecrawl), [Exa MCP](https://github.com/exa-labs/exa-mcp-server), [Tavily MCP](https://github.com/tavily-ai/tavily-mcp), [Jina Reader](https://github.com/jina-ai/reader), [Context7](https://github.com/upstash/context7), [Cloudflare Code Mode](https://blog.cloudflare.com/code-mode/), [Daytona on GitHub](https://github.com/daytonaio/daytona), [Cognee](https://github.com/topoteretes/cognee), [SpecStory](https://github.com/specstoryai/getspecstory), [SpecStory docs](https://docs.specstory.com/integrations/claude-code), [Entire CLI](https://github.com/entireio/cli), [Entire launch post](https://entire.io/blog/hello-entire-world/), [cass](https://github.com/Dicklesworthstone/coding_agent_session_search), [Claude Code History Viewer](https://github.com/jhlee0409/claude-code-history-viewer), [ccusage](https://github.com/ccusage/ccusage), [Langfuse joins ClickHouse](https://langfuse.com/blog/joining-clickhouse), [Langfuse coding agent tracing](https://langfuse.com/resources/engineering/coding-agent-tracing), [Claude Code MCP](https://code.claude.com/docs/en/mcp), [Claude Code on the web](https://code.claude.com/docs/en/claude-code-on-the-web), [Claude Code output styles](https://code.claude.com/docs/en/output-styles), [Claude apps gateway](https://code.claude.com/docs/en/claude-apps-gateway), and our own pages [Context Mode Engine](https://context-mode.com/context-saving) and [Context Mode Insight](https://context-mode.com/insight)
- LLM gateways: [LiteLLM on GitHub](https://github.com/BerriAI/litellm), [LiteLLM tool permission](https://docs.litellm.ai/docs/proxy/guardrails/tool_permission), [LiteLLM plugin marketplace](https://docs.litellm.ai/docs/tutorials/claude_code_plugin_marketplace), [LiteLLM fallbacks](https://docs.litellm.ai/docs/proxy/reliability), [Portkey Claude Code](https://portkey.ai/docs/integrations/libraries/claude-code), [Portkey Codex](https://portkey.ai/docs/integrations/libraries/codex), [Portkey guardrails](https://portkey.ai/docs/product/guardrails), [Portkey gateway on GitHub](https://github.com/Portkey-AI/gateway), [Palo Alto Networks and Portkey](https://www.paloaltonetworks.com/company/press/2026/palo-alto-networks-completes-acquisition-of-portkey-to-secure-ai-agents), [Helicone on GitHub](https://github.com/Helicone/helicone), [Helicone joins Mintlify](https://www.helicone.ai/blog/joining-mintlify), [Helicone Claude Code](https://docs.helicone.ai/integrations/anthropic/claude-code), [Helicone sessions](https://docs.helicone.ai/features/sessions), [OpenRouter Claude Code](https://openrouter.ai/docs/guides/guides/claude-code-integration), [OpenRouter transforms](https://openrouter.ai/docs/features/message-transforms), [Cloudflare AI Gateway](https://developers.cloudflare.com/ai-gateway/), [Cloudflare guardrails](https://developers.cloudflare.com/ai-gateway/features/guardrails/), [Kong AI Gateway](https://developer.konghq.com/ai-gateway/), [Kong AI CLIs](https://developer.konghq.com/ai-gateway/ai-clis/), [Kong prompt compressor](https://developer.konghq.com/plugins/ai-prompt-compressor/), [Vercel AI Gateway](https://vercel.com/docs/ai-gateway), [Vercel coding agents](https://vercel.com/docs/ai-gateway/coding-agents), [Claude apps gateway](https://code.claude.com/docs/en/claude-apps-gateway), [Claude Code LLM gateway](https://code.claude.com/docs/en/llm-gateway), [Bifrost on GitHub](https://github.com/maximhq/bifrost), [Bifrost CLI agents](https://docs.getbifrost.ai/cli-agents/overview), [Bifrost Edge](https://docs.getbifrost.ai/edge/overview), [agentgateway](https://github.com/agentgateway/agentgateway), [claude-code-router](https://github.com/musistudio/claude-code-router), [TensorZero](https://github.com/tensorzero/tensorzero)
- Token and context optimizers: [RTK](https://github.com/rtk-ai/rtk), [RTK site](https://www.rtk-ai.app/), [RTK Pro](https://pro.rtk-ai.app/), [Headroom](https://github.com/headroomlabs-ai/headroom), [Caveman](https://github.com/JuliusBrussee/caveman), [Compresr Context Gateway](https://github.com/Compresr-ai/Context-Gateway), [LLMLingua](https://github.com/microsoft/LLMLingua), [Augment Context Engine](https://www.augmentcode.com/context-engine), [Augment MCP](https://docs.augmentcode.com/context-services/mcp/overview), [Claude Context](https://github.com/zilliztech/claude-context), [Serena](https://github.com/oraios/serena), [Repomix](https://github.com/yamadashy/repomix)
- Agent security and MCP gateways: [Zenity](https://www.zenity.io/use-cases/agent-type/coding-personal-agents), [Lasso](https://www.lasso.security/use-cases/ai-coding-assistants), [Lasso claude-hooks](https://github.com/lasso-security/claude-hooks), [Lasso mcp-gateway](https://github.com/lasso-security/mcp-gateway), [Knostic Kirin](https://www.knostic.ai/ai-coding-security-solution-kirin), [Prisma AIRS](https://www.paloaltonetworks.com/prisma/prisma-ai-runtime-security), [SentinelOne and Prompt Security](https://www.sentinelone.com/press/sentinelone-to-acquire-prompt-security/), [Lakera Guard](https://docs.lakera.ai/docs/quickstart), [Operant AI](https://operant.ai), [MintMCP](https://mintmcp.com), [Obot](https://obot.ai), [Obot on GitHub](https://github.com/obot-platform/obot), [Docker MCP Gateway](https://docs.docker.com/ai/mcp-catalog-and-toolkit/mcp-gateway/), [Docker MCP Gateway on GitHub](https://github.com/docker/mcp-gateway), [Snyk Agent Scan](https://github.com/snyk/agent-scan), [Cisco mcp-scanner](https://github.com/cisco-ai-defense/mcp-scanner), [Cisco DefenseClaw](https://github.com/cisco-ai-defense/defenseclaw), [Invariant Gateway](https://github.com/invariantlabs-ai/invariant-gateway)
- Sandboxes: [Claude Code sandboxing](https://code.claude.com/docs/en/sandboxing), [sandbox-runtime](https://github.com/anthropics/sandbox-runtime), [Claude Code dev container](https://code.claude.com/docs/en/devcontainer), [Codex sandboxing](https://developers.openai.com/codex/concepts/sandboxing.md), [Codex sandboxing (ChatGPT docs)](https://learn.chatgpt.com/docs/sandboxing), [Cursor run modes](https://cursor.com/docs/agent/security/run-modes), [Cursor security](https://cursor.com/docs/agent/security), [Docker Sandboxes](https://docs.docker.com/ai/sandboxes/), [Docker Sandboxes security](https://docs.docker.com/ai/sandboxes/security/), [E2B](https://github.com/e2b-dev/E2B), [Daytona](https://www.daytona.io/docs/en/), [Daytona on GitHub](https://github.com/daytonaio/daytona), [Modal Sandboxes](https://modal.com/docs/guide/sandboxes), [Modal networking](https://modal.com/docs/guide/sandbox-networking)
- Memory layers: [Claude Code memory](https://code.claude.com/docs/en/memory), [Codex memories](https://developers.openai.com/codex/memories), [Anthropic memory tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/memory-tool), [mem0](https://github.com/mem0ai/mem0), [OpenMemory](https://mem0.ai/openmemory), [Graphiti](https://github.com/getzep/graphiti), [supermemory Claude Code](https://supermemory.ai/docs/integrations/claude-code), [supermemory on GitHub](https://github.com/supermemoryai/supermemory), [claude-mem](https://github.com/thedotmack/claude-mem), [Basic Memory](https://github.com/basicmachines-co/basic-memory), [Letta Code memory](https://docs.letta.com/letta-code/memory), [Letta Code on GitHub](https://github.com/letta-ai/letta-code)
- Observability and team analytics: [Langfuse docs](https://langfuse.com/docs), [Langfuse on GitHub](https://github.com/langfuse/langfuse), [Langfuse Claude Code](https://langfuse.com/integrations/other/claude-code), [Langfuse coding agents](https://langfuse.com/coding-agents), [Langfuse coding agent tracing](https://langfuse.com/resources/engineering/coding-agent-tracing), [LangSmith Claude Code](https://docs.langchain.com/langsmith/trace-claude-code), [Phoenix on GitHub](https://github.com/Arize-ai/phoenix), [Phoenix Claude Code](https://arize.com/docs/phoenix/integrations/coding-agents/claude-code), [Braintrust Claude Code](https://www.braintrust.dev/docs/integrations/sdk-integrations/claude-code), [Braintrust Gateway](https://www.braintrust.dev/docs/deploy/gateway), [Datadog Agent Console](https://docs.datadoghq.com/ai_agents_console/), [Datadog LLM Observability](https://docs.datadoghq.com/llm_observability/), [Jellyfish](https://jellyfish.co/platform/ai-impact/), [LinearB](https://linearb.helpdocs.io/article/p4lfzfnu0w-ai-tools-integrations), [DX](https://getdx.com/ai-measurement/), [Swarmia](https://www.swarmia.com/product/ai-impact/), [Faros AI](https://www.faros.ai/)
- Native client features: [Claude Code costs](https://code.claude.com/docs/en/costs), [Claude Code MCP](https://code.claude.com/docs/en/mcp), [Claude Code permissions](https://code.claude.com/docs/en/permissions), [Claude Code hooks](https://code.claude.com/docs/en/hooks), [Claude Code workflows](https://code.claude.com/docs/en/common-workflows.md), [Claude Code analytics](https://code.claude.com/docs/en/analytics), [Claude Code OpenTelemetry](https://code.claude.com/docs/en/monitoring-usage), [Codex config reference](https://developers.openai.com/codex/config-reference), [Codex slash commands](https://developers.openai.com/codex/cli/slash-commands.md), [Cursor analytics](https://cursor.com/docs/account/teams/analytics), [Cursor Admin API](https://cursor.com/docs/account/teams/admin-api), [Anthropic context editing](https://platform.claude.com/docs/en/build-with-claude/context-editing), [Anthropic compaction](https://platform.claude.com/docs/en/build-with-claude/compaction), [Anthropic tool search](https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-search-tool), [Anthropic programmatic tool calling](https://platform.claude.com/docs/en/agents-and-tools/tool-use/programmatic-tool-calling), [Anthropic prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching), [OpenAI prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching)
- Skills and rules: [Claude Code skills](https://code.claude.com/docs/en/skills), [Claude Code marketplaces](https://code.claude.com/docs/en/plugin-marketplaces), [Claude Code plugins for an org](https://code.claude.com/docs/en/plugins/org), [Codex AGENTS.md](https://learn.chatgpt.com/docs/agent-configuration/agents-md), [Codex skills](https://learn.chatgpt.com/docs/build-skills), [Cursor rules](https://cursor.com/docs/context/rules), [Cline rules](https://docs.cline.bot/features/cline-rules), [Roo Code rules](https://roocodeinc.github.io/Roo-Code/features/custom-instructions), [Roo Code on GitHub](https://github.com/RooCodeInc/Roo-Code), [Continue rules](https://docs.continue.dev/customize/deep-dives/rules), [Continue on GitHub](https://github.com/continuedev/continue), [Agent Skills](https://agentskills.io), [skills.sh](https://skills.sh), [anthropics/skills](https://github.com/anthropics/skills)

Something here out of date? Tell us and we fix it.

[Connect your agent](https://context-mode.com/docs/quick-start) [What it is](https://context-mode.com/docs) [See the benchmarks](https://context-mode.com/docs/benchmarks)
