Gateway · Recall for Claude Code and Codex

# Never compact. Never forget. *Send less per request.*

**Recall lets one Claude Code or Codex session run all day without compacting: your client keeps every word, and the model gets your recent work plus an index of the rest.** In a paired run on a 200K window, a long session cost [37.6%](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) less than Claude Code direct and answered [9 of 10](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) questions about early details exactly, against [4 of 10](https://context-mode.com/benchmarks/recall/200k-r2/ab.json).

[37.6%lower cost per long session, 200K window](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) [9 of 10exact answers on early details, 4 of 10 direct](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) [Nevercompacted through the gateway, 3 times direct](https://context-mode.com/benchmarks/recall/200k-r2/ab.json)

Paired run: 3 pre-registered pairs on Claude Opus 5.5, 2026-10-04 ([result JSON](https://context-mode.com/benchmarks/recall/200k-r2/ab.json)). Live figures: the owner's own Claude Code sessions and probes, 2026-10-02. Evidence: [session JSON](https://context-mode.com/benchmarks/data/recall-live-canary-2026-10-02.json) · [probe run 1](https://context-mode.com/benchmarks/data/recall-probe-2026-10-02-run1.json) · [probe run 2](https://context-mode.com/benchmarks/data/recall-probe-2026-10-02-run2.json) · [Method](https://context-mode.com/docs/recall#evaluation)

[Connect your agent →](https://context-mode.com/docs/quick-start) [Read the docs](https://context-mode.com/docs/recall)

## Key facts

What it is

A Context Mode Gateway feature that replaces compaction: older work is archived word for word and indexed, and the agent looks it up when it needs it.

Who it is for

Developers who keep one Claude Code or Codex session open for hours, and teams who pay for those sessions.

Clients

Recall works in Claude Code and Codex. On Codex, a later turn reading back an exact early detail from the archive is not proven yet. In every plan.

Measured result

Long session, 200K window, 3 pairs: [$3.04](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) through the gateway against [$4.86](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) direct, [37.6%](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) lower, [9 of 10](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) exact answers against [4 of 10](https://context-mode.com/benchmarks/recall/200k-r2/ab.json), and direct [compacted 3 times](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) while the gateway session [never compacted](https://context-mode.com/benchmarks/recall/200k-r2/ab.json). On the 1M window: [59.0%](https://context-mode.com/benchmarks/long-session/amendment-4/ab.json) lower, with [10, 10, 10 and 8 of 10](https://context-mode.com/benchmarks/long-session/amendment-4/ab.json) memory answers against 10 of 10 direct. One live session, 2026-10-02 (15:25 to 18:52 UTC): 757,803 tokens per request before, 150,638 where Recall held, 0 upstream errors ([JSON](https://context-mode.com/benchmarks/data/recall-live-canary-2026-10-02.json)). 12 of 12 exact recalls in 2 probe runs ([JSON](https://context-mode.com/benchmarks/data/recall-probe-2026-10-02-run1.json), [JSON](https://context-mode.com/benchmarks/data/recall-probe-2026-10-02-run2.json)).

Limits

Cost measured in Claude Code on Claude Opus 5.5; Codex cost is not measured yet. The agent searches to see older work.

The problem

## Four hours in, Claude Code compacts. *The details go first.*

Near the window, Claude Code swaps your history for a summary. The exact failing line, the rule you gave at the start and the screenshot you pasted are not in it. Until then, every request sends the whole history again, and after each pause the cache is written again at the higher price.

| Four hours in | With Recall | Auto-compact |
| --- | --- | --- |
| "Keep the refund window at 14 days" | Word for word, and settled: the agent does not ask again | In the summary, if it was kept |
| The failing line from test run 3 | Fetched from the archive by id | Gone; the tests run again |
| The error screenshot | Returned as the image | Gone |
| A file read, edited since | Marked "changed at 15:30", newer copy named | Gone |

An example session, to show the mechanism. The measured figures are above and in the [docs](https://context-mode.com/docs/recall#evaluation).

How it works

### Recent work, an index, an archive

## Your client keeps everything. *The model gets what it needs now.*

recent work your newest messages, unchanged: 100K, 150K or 300K tokens, you pick

the index every word you typed, in full, and one line per older step with its id

the archive every older message and tool call, word for word, confirmed before it leaves the model's view

when it acts after a pause, the first time, at the window, or when growth has paid for one cache write

otherwise nothing changes, so the provider keeps reading the prompt from its cache

Finding things

### Exact, and aware of time

## The agent looks it up *instead of asking you.*

by id the exact archived copy; an image comes back as the image

by words this worktree first, then the project or everything when asked

as of what was known at a given time

changed since an old file read says when the file changed and names the newer copy

decisions what you decided stays settled in this worktree; a newer decision replaces the older one

In the decisions check the model did not ask a decided question again, and another worktree of the same project did not see the decision.

In your terminal

## You see when it acts.

```
⚡ context-mode · picked up after a 2 h pause, nothing compacted
⚡ context-mode · recalled 2 items from 2 days ago (a decision, src/billing/quota.ts)
```

Each line shows once. The console lists every lookup under "Your agent looked back" and every decision under "Things you told your agent".

Try it in 5 minutes

## A sub-agent read the log. *You still get the exact line.*

Two tests broke after a teammate's change. A sub-agent ran the suite and sent back only the counts. You ask for the exact failing line, and nothing runs again.

**Setup**

```
npx @context-mode/cli      # sign in and connect Claude Code, then restart Claude Code
git clone https://github.com/expressjs/express.git && cd express && npm install
sed -i.bak "s/'Authentication failed, please check your '/'Login failed, please check your '/" examples/auth/index.js
git commit -qam 'copy: friendlier login error'
claude
```

**Prompts**, in one session:

```
1. Use a subagent to run the full test suite (npm test). I only need the pass and fail counts back from it.
2. Which assertion failed in the auth tests? Quote the exact error line. Don't run the tests again.
```

**What you will see** at prompt 2:

```
test/acceptance/auth.js:36:10
to match /Authentication failed/
...
⏺ context-mode · recalled 10 items from 1 minute ago (a fact, Bash)
```

Run 5 times on 2026-10-04 in real Claude Code sessions: the exact line came back in 5 of 5 runs, and in 5 of 5 new sessions. Output stays in the worktree where it ran, so a second worktree of the same repo starts clean. What each step does is in the [docs](https://context-mode.com/docs/recall#try-it).

Compared

## Next to what you may already use.

|  | Context Mode Recall | Claude Code auto-compact | Memory tools |
| --- | --- | --- | --- |
| Keeps every detail recallable | Yes, all of it | No, a summary | What the tool chose to keep |
| Exact copy | Yes, by id | No | Not stated |
| Knows what changed | Yes: "changed at", "as of" | No | Not stated |
| Cost on a long session | Recent work and an index per request | Grows to the window, then the summary | Full history, plus the tool's own tokens |
| No change to how you work | Yes, after one setup command | Yes, built in | A plugin or hooks on each machine |

Read 2026-10-02: [Claude Code](https://code.claude.com/docs/en/how-claude-code-works), [mem0](https://docs.mem0.ai/integrations/claude-code), [supermemory](https://github.com/supermemoryai/claude-supermemory), [claude-mem](https://github.com/thedotmack/claude-mem). [Full comparison →](https://context-mode.com/docs/recall#compared)

Paired runs

## A long session costs less, *and keeps its details.*

The same 30-step Claude Code session on Claude Opus 5.5, once direct and once through the Gateway with Recall on: 20 work steps, then 10 questions about details from early in the session. The rules were committed before the first pair.

[200K window: cost per session $3.04 through the Gateway, against $4.86 direct: 37.6% lower, cheaper in 3 of 3 pairs](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) [200K window: exact answers 9 of 10 through the Gateway, against 4 of 10 direct](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) [200K window: compactions Never compacted through the Gateway. Direct compacted 3 times](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) [1M window: cost per session $2.97 through the Gateway, against $7.24 direct: 59.0% lower, 4 of 4 pairs](https://context-mode.com/benchmarks/long-session/amendment-4/ab.json)

**On a 200K window.** 3 pairs on 2026-10-04, with the account's Context limit at 100K. Direct [compacted 3 times](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) and lost the values from deleted reports; through the Gateway the session [never compacted](https://context-mode.com/benchmarks/recall/200k-r2/ab.json), Recall [re-laid out 2 times](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) at most in a pair, and those values came back from the archive. Task checks [8 of 8](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) in both arms. [Details →](https://context-mode.com/docs/recall#ab-200k-r2)

**On the 1M window.** 4 pairs on 2026-10-03. Task checks [8 of 8](https://context-mode.com/benchmarks/long-session/amendment-4/ab.json) in both arms, and no compaction in any run. [Details →](https://context-mode.com/docs/recall#ab-result)

Every file: [200K evidence index](https://context-mode.com/benchmarks/recall/index.json) · [1M evidence index](https://context-mode.com/benchmarks/long-session/index.json).

## FAQ

### Does it compact my conversation?

No. Your client keeps its full history. Older work leaves the model's view, not the archive.

### Is anything deleted?

Not within your retention period, 90 days by default. Secrets and e-mail addresses are masked before storage and cannot be read back.

### Does it work with Codex?

Yes. Recall works in Claude Code and Codex. On Codex, a later turn reading back an exact early detail from the archive is not proven yet.

### Does it cost less?

Yes, on a long session: [37.6%](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) less than Claude Code direct on a 200K window, and [59.0%](https://context-mode.com/benchmarks/long-session/amendment-4/ab.json) less on the 1M window, in pre-registered paired runs. Each request carries fewer tokens, and a new layout costs one cache write, made only when it pays back.

### Can I turn it off?

Yes. Set "Recent work in view" to Off in Settings. "Forget this project" deletes its archive.

## One session, *all day.*

*npx @context-mode/cli* connects Claude Code and Codex to the gateway. Recall keeps the evidence; [Memory](https://context-mode.com/memory) keeps the facts it proves.

[Connect your agent →](https://context-mode.com/docs/quick-start) [The technical report →](https://context-mode.com/docs/recall)
