How we prove Context Mode does not break your coding agent

Engineering6 min read

If you put anything between your coding agent and the model, you need to know it will not break a turn, turn off a feature you pay for, or fail when a new model ships. We test Context Mode with the real Claude Code and Codex clients, read the result from each client's own files, and check 31 features in 16 account states for each agent. Here is how it works and what it found.

The problem: a layer you cannot see into

You run a coding agent all day. You add a tool in front of the model to save tokens, add memory or block risky commands. Now you have new worries:

These worries are fair. Claude Code speaks the Anthropic Messages API. Codex speaks the OpenAI Responses API. The tools have different names, reasoning comes back in different blocks, and the two clients write different session files. A change that is correct for one client can break the other.

The short version

We test the real clients, not a mock

The test harness drives the real claude -p and codex exec. Each client gets one small adapter. The adapter does three jobs:

  1. It gives each session a clean start. A fresh config folder for every Claude Code and every Codex session, so one test cannot leak into the next. For a ChatGPT login, the adapter links the login file instead of copying it, so a test never signs your real login out.
  2. It maps tool names to plain kinds. Bash and exec_command are both "shell". Edit and apply_patch are both "write". A scenario asks "was there a shell call?" and does not care which client made it.
  3. It reads the client's own record. For Claude Code, that is its session transcript, sub-agents included. For Codex, it is the rollout file.

Each scenario is written once and runs for every agent. If an agent cannot do a scenario, the matrix shows a red cell. It never hides a row.

We do not trust what the model says

Every scenario checks two witnesses:

The model's reply text is never the proof. The tests use random values the model cannot guess. One scenario writes a random value into in.txt and asks the agent to copy it into out.txt with -OK added. The test then reads out.txt from disk.

Another asks a counting question whose answer is 191, then runs a second turn. That second turn matters: a bug with signed reasoning blocks shows up one turn late, so a one-turn test would miss it.

The 21 scenarios cover plain turns, tool calls, compaction, sub-agents, Memory, Recall, Skills, Protect and CLI setup. The 2026-10-08 run used 74 Claude Code turns and 49 Codex turns.

The matrix also shows fixes as they land. One scenario, upstream-error, stayed red until the fix that cleans up error bodies from the provider went live. Its newest Codex result, from 20:51 UTC, passes 6 of 6 checks.

Your plan state never breaks your agent

A product with plans has a second risk: a feature that is on when it should be off, or off when you paid for it. So we test every feature in every account state, for both agents.

Account stateWhat must happen
ProEvery feature on
Free with requests left (including at 80% and 95%)Every feature on
Free, requests used upFeatures off; requests go straight to the model, unchanged
Pro, card declined, Stripe still retryingEvery feature stays on
Cancelled, refunded, disputed, team memberEach has its own expected result, written down in the matrix

Each state is set up the way production reaches it: with the same plan updates Stripe sends and the same request counts real traffic sends.

In a used-up state, the test checks that the request the gateway forwards is byte for byte the request your agent sent, and that the reply is the provider's reply plus one short notice. Running out of free requests never breaks your agent.

Results from 2026-10-09

AgentCellsPassCorrectly offFailNot applicable
Claude Code49629520100
Codex496286194016

The 16 "not applicable" cells are one feature, Active Context, which works on Claude Code only. This fast layer runs against the shipped gateway code with a fake model, so it spends nothing. It checked all 496 cells per agent in 22 seconds, and it runs inside our normal test suite.

A second, live layer samples cells through the real clients on the production gateway. We have not published a live-layer report yet. We will when one exists.

A new feature cannot skip the check

The matrix is only useful if it is complete. One test scans the shipped code for anything a feature leaves behind: a timing mark, a usage counter, a route, a file the CLI writes. If no one has declared which account states get that feature, the test fails with one fixed message:

declare its entitlement in src/billing/feature-registry.ts

So no feature ships until someone writes down who gets it.

What the tests found

Protect checks each tool call your agent makes and blocks risky ones. Its test suite has an API layer, a 72-case matrix and live turns through both clients.

On 2026-10-08 it found a high-severity problem. Any credential could switch off a Protect rule. A device key lives in the client config on your disk, where the agent itself can read it. In the live run, a key turned off the rule "Delete the root or home folder", and rm -rf ~ then ran in the test.

Now, a change that loosens a rule needs a signed-in person. The agent cannot do it with a key it found on disk. The recorded runs used 28 Claude Code turns and 26 Codex turns.

We test speed with the same care

For the latest rebuild of the request path, unit tests were not enough. We ran two copies of the gateway side by side, old code and new, and sent each request to both at the same moment, across 203 test accounts. The analysis uses 10,000 resamples and 95% intervals.

First we ran the same code on both sides. The first try showed one side 2 to 3 times slower than production, because the test driver started conversations production never starts. We fixed the driver before we trusted any result.

Before, p901,138 ms
After, p90905 ms
Before, p50200 ms
After, p50176 ms
Gateway time before the model call, from the side-by-side test. Not yet confirmed on production traffic.

The test setup fakes the model calls, so quality numbers come from offline tests instead. These numbers are not yet confirmed on production traffic. One quality goal was "not met as written", and our report says so. The full story is in Slipstream.

When a new model ships

One command runs every suite, cheapest first: plan states, billing, the agent scenarios and the real-client suites. A red suite does not stop the next one.

The agent tests take one model per agent. A single model id for both agents is refused, because one id cannot be right for Claude Code and Codex at once. A single "pick the model" option for the full run is not released yet; until then, we run a second check of every feature on the new model.

Our rule: a red cell is fixed in the gateway. We never change the test to make it pass.

How to try it

npx @context-mode/cli

The CLI points Claude Code or Codex at Context Mode. Free gives 1,000 requests once, with every feature and no card. Your own Anthropic or OpenAI account pays the model.