# Recall docs: long sessions that never compact

A long Claude Code session fills the context window, and auto-compact then swaps your history for a summary that drops exact lines, early decisions and screenshots. Recall, a Context Mode Gateway feature, keeps every word and sends the model your recent work plus an index of the rest. In a paired run on a 200K window, a long session cost [37.6%](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) less than Claude Code direct, answered [9 of 10](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) questions about early details exactly against [4 of 10](https://context-mode.com/benchmarks/recall/200k-r2/ab.json), and never compacted.

## Key facts

What it is

A long-session layer in Context Mode Gateway: the model sees a bounded window of recent work and an index of older work, and the agent finds anything older in an archive with `context-mode-search`.

Who it is for

Developers who keep one Claude Code or Codex session open for hours and lose details when it compacts, and teams who pay for long sessions.

Clients

Recall works in Claude Code and Codex. On Codex, a later turn reading back an exact early detail from the archive is not proven yet. In every plan: see [Plans and limits](https://context-mode.com/docs/plans-and-limits).

Measured result

A long Claude Code session on Claude Opus 5.5, 200K window, 3 pre-registered pairs on 2026-10-04: [$3.04](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) through the gateway against [$4.86](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) direct, [37.6%](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) lower. Direct [compacted 3 times](https://context-mode.com/benchmarks/recall/200k-r2/ab.json); through the gateway the session [never compacted](https://context-mode.com/benchmarks/recall/200k-r2/ab.json). On the 1M window: [59.0%](https://context-mode.com/benchmarks/long-session/amendment-4/ab.json) lower, with [10, 10, 10 and 8 of 10](https://context-mode.com/benchmarks/long-session/amendment-4/ab.json) memory answers against 10 of 10 direct. One live session (2026-10-02): 757,803 tokens per request before Recall, 150,638 where it held ([JSON](https://context-mode.com/benchmarks/data/recall-live-canary-2026-10-02.json)). [Evaluation](#evaluation).

Limits

Cost measured in Claude Code on Claude Opus 5.5. Codex cost is not measured yet. The agent searches to see older work.

## Summary

A long Claude Code session fills the model's context window. When it is close to full, Claude Code summarizes the older part and drops the details. Until then, every request sends the whole history again, so each request costs more than the one before.

Recall changes what the model gets, not what your client keeps. Your client still holds every message. The gateway sends the model your recent work, plus a short index of everything older. Before an older message leaves the model's view, the gateway archives it word for word. The agent can get it back by its id, or find it by searching. Nothing is summarized and nothing is thrown away.

What that did in a paired run of a long Claude Code session on a 200K window, against Claude Code direct ([method](#ab-200k-r2)):

[Cost per session 37.6% lower $3.04 through the gateway, $4.86 direct, cheaper in 3 of 3 pairs.](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) [Exact answers about early details 9 of 10 through the gateway, 4 of 10 direct.](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) [Compactions Never compacted through the gateway. Direct compacted 3 times.](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) [Task checks 8 of 8 in both arms, every pair.](https://context-mode.com/benchmarks/recall/200k-r2/ab.json)

And on the owner's own main session ([method](#owner-session)):

[Tokens per request to the model 150,638 mean of the 29 requests where Recall held. 757,803 in the 11 before, while the client kept its full history of about 17 MB.](#owner-session) [Upstream errors 0 in the same window. Claude Code did not compact in it.](#owner-session) [Exact answers about deleted tool output 12 of 12 in a real Claude Code probe. The facts were only in the archive.](#recall-probe) [Decided questions asked again 0 in the decisions check. Another worktree did not see the decision.](#decisions-check)

The paired runs were pre-registered: the rules were committed before the first pair. The owner's session and the probes are real sessions, not a benchmark. Every figure links to its evidence in [Evaluation](#evaluation).

## One long session, two endings

An example of the kind of session Recall is for. A developer works on billing all afternoon in one Claude Code session. Early on, they say: "Keep the refund window at 14 days. Do not touch the webhook retry code." The agent reads 40 files, runs the tests 30 times, and gets a screenshot of an error page. Four hours in, the session passes its window.

| Four hours in | Claude Code auto-compact | Claude Code with Recall |
| --- | --- | --- |
| Your rule "refund window 14 days" | In the summary if the summary kept it, in other words | In the index word for word, and listed as a decision the agent must not reopen |
| "Do not touch the webhook retry code" | Can get lost with the other early instructions | In the index word for word |
| The exact failing line from test run 3 | Gone from the conversation. The agent runs the tests again to see it. | One line in the index points to the archived run. The agent fetches it by id. |
| The error screenshot | Gone | Archived as an image. Its id in the index returns it. |
| The quota file it read at 14:10, edited at 15:30 | Gone | The old read comes back marked "changed at 15:30", with the id of the newer copy |
| Next request | Starts from the summary | Carries recent work and the index; the history in your client is untouched |

The table is an illustration of the mechanism, not a measurement. The measured figures are in [Evaluation](#evaluation).

## The problem

Three things go wrong in a session that runs all day.

- **Compaction loses details.** Close to the window, Claude Code replaces older history with a summary. Its own docs say that instructions from early in the conversation can get lost ([How Claude Code works](https://code.claude.com/docs/en/how-claude-code-works), read 2026-10-02). Exact lines of output, the reason behind a decision, and screenshots are the first to go.
- **The agent asks again.** After a summary, the agent no longer sees what you decided. It asks the same question, or makes the other choice.
- **Cost grows with length.** Every request sends the whole history again. After a pause, the provider's prompt cache has expired, and the whole history is written to the cache again at the higher write price. On a session of 757,803 tokens, that is every token, on every pause.

Memory tools help with the second point: they keep chosen facts. They do not shrink what the client sends, and a fact is not the exact line it came from.

## How it works

Recall acts on the request between your client and the model. The client sends its full history, as it always does. The gateway decides what part of it the model gets.

### When it acts

On most requests Recall changes nothing: the request goes out with the same start as the one before, so the provider reads it from its cache. It lays the request out again only at one of these moments.

| Moment | When | Why then |
| --- | --- | --- |
| **After a pause** | Longer than the prompt cache lasts (5 minutes or 1 hour, as the client asks) | The provider writes the whole prompt again anyway. A shorter prompt costs nothing extra to switch to, and is cheaper to write. |
| **First time** | The first request of a thread that Recall has not laid out yet | One cache write, to start the bounded layout. |
| **The window** | The request would pass your [Context limit](https://context-mode.com/docs/context-limit) | It has to change to fit. This is the point where the client would otherwise compact. |
| **The cost rule** | While the cache is warm, the history keeps growing | Recall adds up what the growing part costs to read on each request. When that sum reaches the cost of one cache write of the short layout, it lays out again. So the write is paid back before it happens. |

It also lays out again when your client itself rewrites older messages, because the cache is written again in that case too.

### What the model gets

The older messages leave the model's view only. They stay in your client, and the archive holds an exact copy that the index points to.

From the top:

1. **The system prompt and tools**, as your client sent them.
2. **The memory block**: facts and decisions that are true now (see [Decisions](#decisions) and [Memory](https://context-mode.com/docs/memory)).
3. **The index of older work.** One block of text, the same bytes on every request until the next layout, so it is read from the cache.
4. **Recent work**: your client's own messages from the cut point to the end, unchanged. The cut point holds a fixed amount of recent work, set in [Settings](#settings), and never moves on a request where nothing above happened.

### The index

The index is built by rules, not by a model, so the same history always gives the same index. A short example:

```
Older part of this conversation, archived by context-mode. Nothing was deleted: every line names its exact archived copy.
Call context-mode-search with {"docid":"<id>"} for the full bytes, or with {"q":"..."} to search it.
2026-10-01 21:58 user: run the billing tests and fix what fails, keep the refund window at 14 days
  → Bash "npm test -- src/billing" (context-mode-msg-3f2a…)
  → Edit src/billing/quota.ts (context-mode-msg-9b1c…)
  ← assistant: Two tests failed in quota.ts. The refund window check used … (context-mode-msg-77de…)
2026-10-01 22:04 user: [img: context-mode-img-5c0e…] this is the error screen, fix it
```

- **Your typed words stay in full.** Every word you typed in the older part is in the index, verbatim. Reminders the client adds are not your words and are left out. A pasted image becomes `[img: id]`.
- **One line per tool call**: the tool, its main input (a path, a command, a pattern or a URL, cut at 120 characters with a visible "…"), and the id of the archived call and result.
- **One line per answer**: its first 160 characters and the id of the full answer.
- **A long index folds by day.** Past 256 KB, whole days older than 2 days become one line each, with their ids. Your own words stay in full even in a folded day.

### Archived before it leaves

A message leaves the model's view only after the archive has confirmed it holds a copy. On the first layout of a very long thread, the archive may need more time than one request allows. Then that request goes out as it would have without Recall, the archive keeps filling in the background, and a later request makes the cut. Nothing leaves the model's view on an unconfirmed write.

If any check on the new layout fails, the gateway sends the request as it would have without Recall. It never sends a half-built layout.

## Finding things: context-mode-search

The agent looks things up with the `context-mode-search` tool. The tool's own description tells the model to search before it asks you something you may have decided, before it says it does not remember, and before it runs a command again only to see an old result.

exact copy `{"docid": "<id>"}` returns the archived item as it was stored, paged with `offset` when long. An image id returns the image.

by words `{"q": "refund window"}` searches the archive and your facts.

scope `"scope"`: this worktree by default; `"project"` for every worktree of the repository; `"all"` for everything on your account.

as of `"as_of": "2026-10-01T22:00:00Z"` asks what was known then: only items from before that time, and facts that were true at it.

kinds `"kinds": ["msg", "tool", "img", "fact"]` narrows the result.

Results come back with current facts first, then archived items ranked by how well they match, how close they are to this conversation, and how recent they are. Each item says where and when it happened.

### "Changed since"

An old file read can be out of date. When the archive holds a read of a file and a later edit of the same file in the same worktree, the result says so: `changed at 14:02, newer: <id>`. A git checkout, pull, merge, rebase or reset marks the worktree's older file reads as "may have changed since". The tool tells the model to read the file on disk for its current content, and to prefer the newer item.

### Pointers in your turn

When your new message names a file path or an error text that matches older archived work, the gateway adds one short line naming the archived item and its age. It adds the pointer, never the content, and at most 3 lines per turn.

## Decisions

A decision you made once should not come back. Recall keeps decisions as facts, tied to the moment they came from.

- **How one is captured.** The agent asks a question or offers options, and you answer ("yes", "B", "do it", a number). Or the agent asks, then acts, and your next message does not object. The question and your answer are both kept, with the id of the turn.
- **What the model sees.** A block at the start of the conversation lists the decisions already made in this worktree, newest first, and tells the model they are settled: act on them, do not ask again, do not reopen them unless you do.
- **Newer replaces older.** A new decision on the same question closes the old one. The old one shows as replaced, as history only.
- **Only what you said or allowed.** A decision a model guessed from free text stays searchable, but never enters the "settled" list.
- **Per worktree.** A decision in one worktree is not shown in another. Say "for the whole project" or "always" and it applies wider.

In the console, decisions are listed under "Things you told your agent". Each one opens the moment it came from, and "Forget" closes it.

## Scopes

| Level | Means | Shared by default |
| --- | --- | --- |
| **Worktree** | One topic of work | Yes: the default for search, decisions and pointers |
| Session, in one worktree | One client process | Several sessions in one worktree share the worktree's archive, facts and decisions. The current conversation ranks first. |
| Project | The repository, all its worktrees | Only when the search asks for `"project"` . Each result names its worktree and branch. |
| Account | You | Facts about you, such as your name and how you like to work |

So two worktrees of one project count as two topics: the decisions and pointers of one never show in the other unless you ask. Two sessions in one worktree work as one memory.

Recall follows the Memory scope: a search can only reach as far as the Memory page lets this worktree see, and tool output and files read stay in the worktree they came from.

## Settings

Recent work in view **Lean**: newest 100K tokens. **Balanced** (default): newest 150K. **Wide**: newest 300K. **Off**: Recall does not act. Never more than half your [Context limit](https://context-mode.com/docs/context-limit). Older work is still searchable. Smaller is cheaper per request.

Retention How long the archive keeps older work: 30, 90 or 365 days, or forever. Default 90 days.

Team sharing Share a project's archive with your team. Off by default, and set per project.

Forget this project Deletes the project's saved history and the facts learned from it. Your settings, other projects and usage stay. The confirm step names the project and how many items it deletes.

Balanced is the default because it keeps a full working stretch in view. A smaller window costs less per request.

## What you see in the terminal

When Recall acts, the gateway adds one line to the reply, in the same banner as its other notes:

```
⚡ context-mode · picked up after a 2 h pause, nothing compacted
⚡ context-mode · recalled 2 items from 2 days ago (a decision, src/billing/quota.ts)
```

- **Picked up** shows on the request that laid the session out again after a pause. When the ledger has a saving for that request, the line adds it.
- **Recalled** shows when the agent's search brought back items: how many, the age of the oldest, and at most two labels such as "a decision", a file path or "a screenshot".
- Each line shows once. The same items are not announced twice in one conversation. Subagents print neither line.

In the console, "Your agent looked back" lists each lookup ("the refund decision, found in billing, 2 h ago"), and a long session reads as one line: how many hours your agent remembers, and what each request carried against the full history.

## What you save

The console shows a "saved by Recall" number on each request Recall touched, and the total under Saved. It compares two prices for the same request:

- **What you paid.** The request as Recall sent it, at the price the model provider charged.
- **What you would have paid.** The same request as the gateway sends it without Recall: your whole history, never more than the model's window. We price it warm, as if the cache already held everything the previous request sent and only the new part is written. Only after a pause longer than the cache lasts do we price it cold, because it would be cold without Recall too.

Saved by Recall is the second price minus the first. Only the model's first read of the request is compared; its reply and later calls cost the same either way.

Recall's own costs and failures count against it. A new layout costs one cache write. When Recall fails on a request, or lets go of a session, the full history is written to the cache again, so that request costs more than it would have without Recall. Those costs are subtracted, so the number can be below zero. It counts every main request from the first one Recall touched, so a loss after Recall stops is still Recall's. Subagents are not counted.

## Try it in 5 minutes

A teammate changed a login message and two tests broke. A sub-agent ran the suite and sent back only the counts. Then you ask for the exact failing line, and nothing runs again.

### Setup

```
npx @context-mode/cli      # sign in and connect Claude Code, then restart Claude Code
git clone https://github.com/expressjs/express.git && cd express && npm install
# play the teammate: a copy change to the login error
sed -i.bak "s/'Authentication failed, please check your '/'Login failed, please check your '/" examples/auth/index.js
git commit -qam 'copy: friendlier login error'
claude
```

### Prompts

One session, one prompt at a time:

```
Use a subagent to run the full test suite (npm test). I only need the pass and fail counts back from it.
```

```
Which assertion failed in the auth tests? Quote the exact error line. Don't run the tests again.
```

### What you will see

```
# prompt 1: counts only. The sub-agent read the log; your session did not.

# prompt 2: no test run, the exact lines
test/acceptance/auth.js:36:10
to match /Authentication failed/
      at Test.<anonymous> (test/acceptance/auth.js:36:10)

test/acceptance/auth.js:51:14
to match /Authentication failed/
      at Test.<anonymous> (test/acceptance/auth.js:51:14)

Both assertions expect the response body to contain /Authentication failed/, but the actual response contains "Login failed, please check your username and password..." instead.

⏺ context-mode · recalled 10 items from 1 minute ago (a fact, Bash)
```

- **Prompt 1.** The counts come from the sub-agent. In a plain clone, `npm test` prints 1259 passing and 2 failing. We ran inside a git worktree in a hidden folder, where some of express's static-file tests also fail, so our sub-agents reported 57 failing.
- **Prompt 2.** The error lines were only in the sub-agent's tool output. The "recalled" line shows that the gateway looked them up in the archive. Check them yourself with `npm test`: the same two auth assertions show up, at `test/acceptance/auth.js` lines 36 and 51.
- **In the console.** The Memory page has a "Where recall reaches" card. It shows each worktree and how many items it keeps, for example "keeps 9 here".

### Two worktrees of one repo

Now open a second worktree of the same repo, for example with `claude --worktree router-cache`, and ask:

```
Paste the exact error from the auth test run earlier today. Don't run the tests.
```

The agent says it has no record of that run. Tool output, files read and logs stay in the worktree where they happened. Another branch may have different code, so an old log from there could mislead the agent.

### The next session

Close the session, start a new one in the same folder, and ask:

```
Paste the exact error from the auth test run before lunch, I need it for the bug ticket. Don't run the tests again.
```

On the first message of a new session, the gateway searches the archive itself and puts the matching lines of the earlier run in front of the agent. The agent quotes the two failing assertions from the earlier run without running anything: 5 of 5 runs.

- **Decisions travel through Memory.** A decision you type ("a plain Map, not lru-cache, no new dependency") reaches your next session in the same folder through Memory, in Claude Code and in Codex. See [Memory](https://context-mode.com/docs/memory#try-it).

We ran this five times on 2026-10-04 in real Claude Code sessions (`claude -p` 2.1.283, Claude Haiku 4.5) through the gateway, each on a new test account. Prompt 2 quoted the exact line in 5 of 5 runs, and so did the next session. Recall is on for every account by default; you turn it off on the Memory page.

**Why it matters.** The exact line you need is still there hours later, even when a sub-agent read it. You do not run a slow suite again just to see one line.

## Recall and Memory together

Recall and [Memory](https://context-mode.com/docs/memory) are two halves of one memory.

Memory The current truth: short facts with two clocks, when each was true and when the gateway learned it.

Recall The evidence: every message and tool call as it happened, with when, where and who.

Each fact links to the moment it came from. When a fact changes, the old one is closed, and the archived moment that set it stays. So the agent can answer "what is true now" from Memory, and "where did that come from" or "what did the file say then" from Recall.

## Evaluation

Two pre-registered paired runs of a long Claude Code session, an offline replay, live checks and the owner's own session. Every number links to the file it comes from.

### Long session on a 200K window: 37.6% lower cost, 9 of 10 exact answers

- **Cost.** [$3.04](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) per session through the gateway against [$4.86](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) direct, [37.6%](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) lower ([30.7 to 44.3%](https://context-mode.com/benchmarks/recall/200k-r2/ab.json)), cheaper in [3 of 3](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) pairs.
- **Quality.** [9 of 10](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) exact answers about early details through the gateway, [4 of 10](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) direct. Direct [compacted 3 times](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) and lost the values from deleted reports; through the gateway the session [never compacted](https://context-mode.com/benchmarks/recall/200k-r2/ab.json), and those values came back from the archive. Task checks [8 of 8](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) in both arms.
- **Recall at work.** Recall acted in every pair: it [re-laid out 2 times](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) at most in one pair. Its thread read took [5 ms](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) at p50.
- **Method.** 3 pairs on 2026-10-04. Claude Code 2.1.283 on Claude Opus 5.5 with its real 200K window, the gateway account's [Context limit](https://context-mode.com/docs/context-limit) at 100K. Each pair ran the same 30 steps direct and through the gateway: 20 work steps, then 10 questions about early details. The rules were committed before the first pair. Evidence: [result JSON](https://context-mode.com/benchmarks/recall/200k-r2/ab.json), [the plan](https://context-mode.com/benchmarks/prereg/recall-200k-r2-prereg.md), [every file](https://context-mode.com/benchmarks/recall/index.json).

### Long session on the 1M window: 59.0% lower cost

- **Cost.** [$2.97](https://context-mode.com/benchmarks/long-session/amendment-4/ab.json) per session through the gateway against [$7.24](https://context-mode.com/benchmarks/long-session/amendment-4/ab.json) direct, [59.0%](https://context-mode.com/benchmarks/long-session/amendment-4/ab.json) lower (55.6 to 61.3%), lower in [4 of 4](https://context-mode.com/benchmarks/long-session/amendment-4/ab.json) pairs.
- **Quality.** Task checks [8 of 8](https://context-mode.com/benchmarks/long-session/amendment-4/ab.json) in both arms. No compaction, re-ask or upstream 4xx error in any of the 8 runs. Exact answers to the 10 memory questions: [10, 10, 10 and 8 of 10](https://context-mode.com/benchmarks/long-session/amendment-4/ab.json) through the gateway, 10 of 10 direct in every pair.
- **Method.** 4 pairs on 2026-10-03. Claude Code 2.1.283 on Claude Opus 5.5, its real 1M window, thinking off. The same 30 steps direct and through the gateway with Recall on; on the questions, the gateway arm used the archive search. The rules were committed before the first pair. Evidence: [result JSON](https://context-mode.com/benchmarks/long-session/amendment-4/ab.json), [the plan](https://context-mode.com/benchmarks/prereg/long-session-amendment-4-prereg.md), [every file](https://context-mode.com/benchmarks/long-session/index.json).

**What it means for you.** A long session costs less than it does direct, and an early detail comes back from the archive, not from a summary.

### Quality replay, no model

- **Method.** The 1M-window session above, rebuilt and run through Recall's real code, request by request. No model call and no network, so it cost $0. "In the body" means the exact answer is in what the gateway sends to the model.
- **The index and auto-recall.** Answers in the body: [5 of 10](https://context-mode.com/benchmarks/recall/replay/visibility.json) with the old index, [8 of 10](https://context-mode.com/benchmarks/recall/replay/visibility.json) with the new index, [10 of 10](https://context-mode.com/benchmarks/recall/replay/visibility.json) with auto-recall, which puts the matching archived line in before the question.
- **The paired run's questions.** [17 of 17](https://context-mode.com/benchmarks/recall/replay/pairs.json) replayed questions are in the body under the new rules, with [0 re-layouts](https://context-mode.com/benchmarks/recall/replay/pairs.json).
- **A longer session on a 200K client.** Recall acts [4 times](https://context-mode.com/benchmarks/recall/replay/pairs.json), [0 requests](https://context-mode.com/benchmarks/recall/replay/pairs.json) go over 200K, and [10 of 10](https://context-mode.com/benchmarks/recall/replay/pairs.json) answers stay in the body. The old rules sent [3 requests](https://context-mode.com/benchmarks/recall/replay/pairs.json) over 200K and had [4 of 10](https://context-mode.com/benchmarks/recall/replay/pairs.json).

### Live checks

- **Claude Code.** A real `claude -p` turn on the owner's account: [passed](https://context-mode.com/benchmarks/recall/live/claude-code-check.json). The second turn showed the "recalled" line, and its answer held a value only the archive had.
- **Codex.** [2 of 2](https://context-mode.com/benchmarks/recall/live/codex-check.json) runs completed, and the answer held the value.

### Read speed

- **Objects beside the gateway.** New thread objects now open in the gateway's own region: [4 of 5](https://context-mode.com/benchmarks/recall/speed/placement.json) did on the first try. Recall's thread read, p50 / p90: [18 / 297 ms](https://context-mode.com/benchmarks/recall/speed/placement.json) before, [16 / 54 ms](https://context-mode.com/benchmarks/recall/speed/placement.json) after.
- **Under load.** Bursts of 10 sub-agents every 75 seconds: thread read p50 [15 ms](https://context-mode.com/benchmarks/recall/speed/load-rerun.json), p99 [145 ms](https://context-mode.com/benchmarks/recall/speed/load-rerun.json), with [0 read failures](https://context-mode.com/benchmarks/recall/speed/load-rerun.json) in every window.

### The owner's main session

- **Session.** The owner's main Claude Code session, 2026-10-02, from 15:25 to 18:52 UTC (Recall turned on at 17:25). Read from the gateway's ledger.
- **Before.** 11 requests, 757,803 tokens per request on average.
- **After.** 29 requests where Recall held the layout: 150,638 tokens per request on average. On 6 requests Recall stepped aside and sent the full history, 799,413 tokens. The client still held its full history of about 17 MB.
- **Health.** 0 upstream errors. Claude Code did not compact in this window.
- **Evidence.** [Session JSON](https://context-mode.com/benchmarks/data/recall-live-canary-2026-10-02.json).

### Recall probe: deleted tool output

- **Method.** Real `claude -p` on Claude Opus 5.5, through the gateway. Each run planted 6 facts in tool output, then deleted the files from disk, so the facts lived only in the archive. 2 runs: 12 questions with one exact answer each.
- **Result.** 12 of 12 exact answers. The model found every one by searching with words, not by id. Evidence: [run 1 JSON](https://context-mode.com/benchmarks/data/recall-probe-2026-10-02-run1.json), [run 2 JSON](https://context-mode.com/benchmarks/data/recall-probe-2026-10-02-run2.json).

### Decisions check

- **Method.** The user answers a question early in a session. Later, in the same conversation and in a new one, the agent reaches the point where it would ask it again. A second worktree of the same project runs the same task.
- **Result.** 0 re-asks: the model followed the decision. The decision stays with its own worktree.

## Verify it yourself

Every evidence file holds the token counts of each call and the list price we used, so you can recompute every dollar yourself. A checking script is available on request.

## Compared with other approaches

From each vendor's own pages, read on 2026-10-02. "Not stated" means the page we read does not say.

| Approach | Keeps every detail recallable | Exact copy | Knows what changed | Cost on a long session | No change to how you work |
| --- | --- | --- | --- | --- | --- |
| **Context Mode Recall** | Yes. Every older message and tool call is archived, word for word except masked secrets and e-mail addresses. | Yes, by id. Images come back as images. | Yes. A file read later edited says "changed at"; "as of" asks what was known then. | Each request carries your recent work and the index, not the whole history. | Yes, after one setup command. |
| [Claude Code auto-compact](https://code.claude.com/docs/en/how-claude-code-works) | No. Older history becomes a summary; early instructions can get lost. | No | No | Grows with every turn until the window, then drops to the summary. | Yes, it is built in. |
| Memory tools: [mem0](https://docs.mem0.ai/integrations/claude-code) , [supermemory](https://github.com/supermemoryai/claude-supermemory) , [claude-mem](https://github.com/thedotmack/claude-mem) | What the tool chose to keep: facts, notes or saved conversations | Not stated | Not stated | The client still sends its full history and compacts at its window. The tool adds its own tokens. | A plugin or hooks on each machine. |

**Where Context Mode differs.** Recall does not choose what to keep. It keeps everything, points to it from the index, and lets the agent fetch the exact item. It also shortens what each request carries, which a memory tool does not do.

**Where others are stronger.**

- Auto-compact needs no account and no network hop, and it works the same in every Claude Code setup.
- Memory tools cover more clients, such as Cursor and OpenCode. Recall works in Claude Code and Codex only.
- A summary is short. Recall's index grows with the session, and an agent that does not search can miss what only the archive holds.

## Privacy and retention

- **Where it lives.** In your account's own store on the gateway. Your account is the wall between tenants.
- **Masking.** Before archived text is stored, known secret formats and e-mail addresses are masked, as in [Search](https://context-mode.com/docs/search). Everything else reads back word for word.
- **Retention.** Older work is kept for your retention period (default 90 days), then deleted each night. Facts learned from it stay, with their link marked as expired.
- **Deletion.** "Forget this project" deletes the project's archive, its search index and every fact that came from it.
- **Team.** Nothing is shared with your team unless you turn on sharing for that project.
- **Subagents.** A subagent searches in its parent conversation's scope.

## Limits

- **Codex.** Recall works in Claude Code and Codex. On Codex, a later turn reading back an exact early detail from the archive is not proven yet, and its cost is not measured yet.
- **Sample size.** The paired runs have 3 and 4 pairs on Claude Opus 5.5. The token figures come from one person's real sessions. The recall probe has 12 questions; the decisions check is one scenario.
- **The agent must search.** An older detail is in the archive, not in view. If the model does not search, it does not see it. The index and the tool description tell it when to search.
- **A cache write per layout.** Each new layout writes the shorter prompt to the cache once. Recall only does it at the moments listed above.
- **Masked text.** A masked secret or e-mail address cannot be read back.

## FAQ for CTOs

### Does Recall compact my conversation?

No. Your client keeps its full history, and the gateway never summarizes it. The model gets your recent work and an index; everything older is in the archive, word for word.

### Is anything deleted?

Not while it is in your retention period (default 90 days). Older messages leave the model's view, not the archive. Only masked secrets and e-mail addresses cannot be read back.

### Does it work with Codex?

Yes. Recall works in Claude Code and Codex. On Codex, a later turn reading back an exact early detail from the archive is not proven yet.

### Does it cost less?

Yes, on a long session. In the paired runs, a long session through the gateway with Recall on cost [37.6%](https://context-mode.com/benchmarks/recall/200k-r2/ab.json) less than Claude Code direct on a 200K window, and [59.0%](https://context-mode.com/benchmarks/long-session/amendment-4/ab.json) less on the 1M window. Each request carries fewer tokens: 150,638 instead of 757,803 on the owner's session. A new layout costs one cache write, and Recall makes one only after a pause, the first time, at the window, or when the cost rule says it has paid for itself.

### Can I turn it off?

Yes. Set "Recent work in view" to Off in Settings. Your archive stays until its retention ends, and "Forget this project" deletes it now.

### Can I get an old screenshot back?

Yes. A pasted image is archived as an image, and its id in the index returns it.

Compared with other tools: see the [landscape](https://context-mode.com/docs/landscape#memory).

[Connect your agent](https://context-mode.com/docs/quick-start) [Recall overview](https://context-mode.com/recall) [Memory](https://context-mode.com/docs/memory)
