Gateway · Recall for Claude Code and Codex

Never compact. Never forget. Send less per request.

Recall lets one Claude Code or Codex session run all day without compacting: your client keeps every word, and the model gets your recent work plus an index of the rest. In a paired run on a 200K window, a long session cost 37.6% less than Claude Code direct and answered 9 of 10 questions about early details exactly, against 4 of 10.

Paired run: 3 pre-registered pairs on Claude Opus 5.5, 2026-10-04 (result JSON). Live figures: the owner's own Claude Code sessions and probes, 2026-10-02. Evidence: session JSON · probe run 1 · probe run 2 · Method

Key facts

What it is
A Context Mode Gateway feature that replaces compaction: older work is archived word for word and indexed, and the agent looks it up when it needs it.
Who it is for
Developers who keep one Claude Code or Codex session open for hours, and teams who pay for those sessions.
Clients
Recall works in Claude Code and Codex. On Codex, a later turn reading back an exact early detail from the archive is not proven yet. In every plan.
Measured result
Long session, 200K window, 3 pairs: $3.04 through the gateway against $4.86 direct, 37.6% lower, 9 of 10 exact answers against 4 of 10, and direct compacted 3 times while the gateway session never compacted. On the 1M window: 59.0% lower, with 10, 10, 10 and 8 of 10 memory answers against 10 of 10 direct. One live session, 2026-10-02 (15:25 to 18:52 UTC): 757,803 tokens per request before, 150,638 where Recall held, 0 upstream errors (JSON). 12 of 12 exact recalls in 2 probe runs (JSON, JSON).
Limits
Cost measured in Claude Code on Claude Opus 5.5; Codex cost is not measured yet. The agent searches to see older work.

The problem

Four hours in, Claude Code compacts. The details go first.

Near the window, Claude Code swaps your history for a summary. The exact failing line, the rule you gave at the start and the screenshot you pasted are not in it. Until then, every request sends the whole history again, and after each pause the cache is written again at the higher price.

Four hours inWith RecallAuto-compact
"Keep the refund window at 14 days"Word for word, and settled: the agent does not ask againIn the summary, if it was kept
The failing line from test run 3Fetched from the archive by idGone; the tests run again
The error screenshotReturned as the imageGone
A file read, edited sinceMarked "changed at 15:30", newer copy namedGone

An example session, to show the mechanism. The measured figures are above and in the docs.


Your client keeps everything. The model gets what it needs now.

recent workyour newest messages, unchanged: 100K, 150K or 300K tokens, you pick
the indexevery word you typed, in full, and one line per older step with its id
the archiveevery older message and tool call, word for word, confirmed before it leaves the model's view
when it actsafter a pause, the first time, at the window, or when growth has paid for one cache write
otherwisenothing changes, so the provider keeps reading the prompt from its cache

The agent looks it up instead of asking you.

by idthe exact archived copy; an image comes back as the image
by wordsthis worktree first, then the project or everything when asked
as ofwhat was known at a given time
changed sincean old file read says when the file changed and names the newer copy
decisionswhat you decided stays settled in this worktree; a newer decision replaces the older one

In the decisions check the model did not ask a decided question again, and another worktree of the same project did not see the decision.


In your terminal

You see when it acts.

⚡ context-mode · picked up after a 2 h pause, nothing compacted
⚡ context-mode · recalled 2 items from 2 days ago (a decision, src/billing/quota.ts)

Each line shows once. The console lists every lookup under "Your agent looked back" and every decision under "Things you told your agent".


Try it in 5 minutes

A sub-agent read the log. You still get the exact line.

Two tests broke after a teammate's change. A sub-agent ran the suite and sent back only the counts. You ask for the exact failing line, and nothing runs again.

Setup

npx @context-mode/cli      # sign in and connect Claude Code, then restart Claude Code
git clone https://github.com/expressjs/express.git && cd express && npm install
sed -i.bak "s/'Authentication failed, please check your '/'Login failed, please check your '/" examples/auth/index.js
git commit -qam 'copy: friendlier login error'
claude

Prompts, in one session:

1. Use a subagent to run the full test suite (npm test). I only need the pass and fail counts back from it.
2. Which assertion failed in the auth tests? Quote the exact error line. Don't run the tests again.

What you will see at prompt 2:

test/acceptance/auth.js:36:10
to match /Authentication failed/
...
⏺ context-mode · recalled 10 items from 1 minute ago (a fact, Bash)

Run 5 times on 2026-10-04 in real Claude Code sessions: the exact line came back in 5 of 5 runs, and in 5 of 5 new sessions. Output stays in the worktree where it ran, so a second worktree of the same repo starts clean. What each step does is in the docs.


Compared

Next to what you may already use.

Context Mode RecallClaude Code auto-compactMemory tools
Keeps every detail recallableYes, all of itNo, a summaryWhat the tool chose to keep
Exact copyYes, by idNoNot stated
Knows what changedYes: "changed at", "as of"NoNot stated
Cost on a long sessionRecent work and an index per requestGrows to the window, then the summaryFull history, plus the tool's own tokens
No change to how you workYes, after one setup commandYes, built inA plugin or hooks on each machine

Read 2026-10-02: Claude Code, mem0, supermemory, claude-mem. Full comparison →


Paired runs

A long session costs less, and keeps its details.

The same 30-step Claude Code session on Claude Opus 5.5, once direct and once through the Gateway with Recall on: 20 work steps, then 10 questions about details from early in the session. The rules were committed before the first pair.

On a 200K window. 3 pairs on 2026-10-04, with the account's Context limit at 100K. Direct compacted 3 times and lost the values from deleted reports; through the Gateway the session never compacted, Recall re-laid out 2 times at most in a pair, and those values came back from the archive. Task checks 8 of 8 in both arms. Details →

On the 1M window. 4 pairs on 2026-10-03. Task checks 8 of 8 in both arms, and no compaction in any run. Details →

Every file: 200K evidence index · 1M evidence index.


FAQ

Does it compact my conversation?

No. Your client keeps its full history. Older work leaves the model's view, not the archive.

Is anything deleted?

Not within your retention period, 90 days by default. Secrets and e-mail addresses are masked before storage and cannot be read back.

Does it work with Codex?

Yes. Recall works in Claude Code and Codex. On Codex, a later turn reading back an exact early detail from the archive is not proven yet.

Does it cost less?

Yes, on a long session: 37.6% less than Claude Code direct on a 200K window, and 59.0% less on the 1M window, in pre-registered paired runs. Each request carries fewer tokens, and a new layout costs one cache write, made only when it pays back.

Can I turn it off?

Yes. Set "Recent work in view" to Off in Settings. "Forget this project" deletes its archive.

One session, all day.

npx @context-mode/cli connects Claude Code and Codex to the gateway. Recall keeps the evidence; Memory keeps the facts it proves.