How we prove Context Mode does not break your coding agent
If you put anything between your coding agent and the model, you need to know it will not break a turn, turn off a feature you pay for, or fail when a new model ships. We test Context Mode with the real Claude Code and Codex clients, read the result from each client's own files, and check 31 features in 16 account states for each agent. Here is how it works and what it found.
The problem: a layer you cannot see into
You run a coding agent all day. You add a tool in front of the model to save tokens, add memory or block risky commands. Now you have new worries:
- Will it break a turn that worked yesterday?
- Will it work the same for Claude Code and for Codex?
- If my free requests run out, or my card fails, does my agent stop?
- When a new model ships, who checks that it still works?
These worries are fair. Claude Code speaks the Anthropic Messages API. Codex speaks the OpenAI Responses API. The tools have different names, reasoning comes back in different blocks, and the two clients write different session files. A change that is correct for one client can break the other.
The short version
- 21 test scenarios run on both real clients. On 2026-10-08, Claude Code passed 21 of 21 and Codex passed 21 of 21.
- An entitlement matrix checks 31 features × 16 account states × 2 agents: 496 cells per agent, 0 failures.
- When your free requests run out, nothing breaks. Requests go straight to the model, unchanged.
- A payment problem never reduces service. Pro stays on while Stripe retries a declined payment.
- The tests found a high-severity Protect problem, and we fixed it.
We test the real clients, not a mock
The test harness drives the real claude -p and codex exec. Each client gets one small adapter. The adapter does three jobs:
- It gives each session a clean start. A fresh config folder for every Claude Code and every Codex session, so one test cannot leak into the next. For a ChatGPT login, the adapter links the login file instead of copying it, so a test never signs your real login out.
- It maps tool names to plain kinds.
Bashandexec_commandare both "shell".Editandapply_patchare both "write". A scenario asks "was there a shell call?" and does not care which client made it. - It reads the client's own record. For Claude Code, that is its session transcript, sub-agents included. For Codex, it is the rollout file.
Each scenario is written once and runs for every agent. If an agent cannot do a scenario, the matrix shows a red cell. It never hides a row.
We do not trust what the model says
Every scenario checks two witnesses:
- The client side: the transcript, or a file the agent wrote to disk.
- The gateway side: the status, stream and timing of each request, and the account's usage records.
The model's reply text is never the proof. The tests use random values the model cannot guess. One scenario writes a random value into in.txt and asks the agent to copy it into out.txt with -OK added. The test then reads out.txt from disk.
Another asks a counting question whose answer is 191, then runs a second turn. That second turn matters: a bug with signed reasoning blocks shows up one turn late, so a one-turn test would miss it.
The 21 scenarios cover plain turns, tool calls, compaction, sub-agents, Memory, Recall, Skills, Protect and CLI setup. The 2026-10-08 run used 74 Claude Code turns and 49 Codex turns.
The matrix also shows fixes as they land. One scenario, upstream-error, stayed red until the fix that cleans up error bodies from the provider went live. Its newest Codex result, from 20:51 UTC, passes 6 of 6 checks.
Your plan state never breaks your agent
A product with plans has a second risk: a feature that is on when it should be off, or off when you paid for it. So we test every feature in every account state, for both agents.
| Account state | What must happen |
|---|---|
| Pro | Every feature on |
| Free with requests left (including at 80% and 95%) | Every feature on |
| Free, requests used up | Features off; requests go straight to the model, unchanged |
| Pro, card declined, Stripe still retrying | Every feature stays on |
| Cancelled, refunded, disputed, team member | Each has its own expected result, written down in the matrix |
Each state is set up the way production reaches it: with the same plan updates Stripe sends and the same request counts real traffic sends.
In a used-up state, the test checks that the request the gateway forwards is byte for byte the request your agent sent, and that the reply is the provider's reply plus one short notice. Running out of free requests never breaks your agent.
Results from 2026-10-09
| Agent | Cells | Pass | Correctly off | Fail | Not applicable |
|---|---|---|---|---|---|
| Claude Code | 496 | 295 | 201 | 0 | 0 |
| Codex | 496 | 286 | 194 | 0 | 16 |
The 16 "not applicable" cells are one feature, Active Context, which works on Claude Code only. This fast layer runs against the shipped gateway code with a fake model, so it spends nothing. It checked all 496 cells per agent in 22 seconds, and it runs inside our normal test suite.
A second, live layer samples cells through the real clients on the production gateway. We have not published a live-layer report yet. We will when one exists.
A new feature cannot skip the check
The matrix is only useful if it is complete. One test scans the shipped code for anything a feature leaves behind: a timing mark, a usage counter, a route, a file the CLI writes. If no one has declared which account states get that feature, the test fails with one fixed message:
declare its entitlement in src/billing/feature-registry.ts
So no feature ships until someone writes down who gets it.
What the tests found
Protect checks each tool call your agent makes and blocks risky ones. Its test suite has an API layer, a 72-case matrix and live turns through both clients.
On 2026-10-08 it found a high-severity problem. Any credential could switch off a Protect rule. A device key lives in the client config on your disk, where the agent itself can read it. In the live run, a key turned off the rule "Delete the root or home folder", and rm -rf ~ then ran in the test.
Now, a change that loosens a rule needs a signed-in person. The agent cannot do it with a key it found on disk. The recorded runs used 28 Claude Code turns and 26 Codex turns.
We test speed with the same care
For the latest rebuild of the request path, unit tests were not enough. We ran two copies of the gateway side by side, old code and new, and sent each request to both at the same moment, across 203 test accounts. The analysis uses 10,000 resamples and 95% intervals.
First we ran the same code on both sides. The first try showed one side 2 to 3 times slower than production, because the test driver started conversations production never starts. We fixed the driver before we trusted any result.
- The error rate fell from 0.06-0.86% to 0%.
- Load runs at 1x and 1.5x normal traffic, 2 minutes each, had 0 errors and 0 memory kills.
The test setup fakes the model calls, so quality numbers come from offline tests instead. These numbers are not yet confirmed on production traffic. One quality goal was "not met as written", and our report says so. The full story is in Slipstream.
When a new model ships
One command runs every suite, cheapest first: plan states, billing, the agent scenarios and the real-client suites. A red suite does not stop the next one.
The agent tests take one model per agent. A single model id for both agents is refused, because one id cannot be right for Claude Code and Codex at once. A single "pick the model" option for the full run is not released yet; until then, we run a second check of every feature on the new model.
Our rule: a red cell is fixed in the gateway. We never change the test to make it pass.
How to try it
npx @context-mode/cli
The CLI points Claude Code or Codex at Context Mode. Free gives 1,000 requests once, with every feature and no card. Your own Anthropic or OpenAI account pays the model.
- Set up: Quick start
- See how Protect blocks risky tool calls: Protect and Protect docs
- What happens when requests run out: Plans and limits