Recall docs: long sessions that never compact

A long Claude Code session fills the context window, and auto-compact then swaps your history for a summary that drops exact lines, early decisions and screenshots. Recall, a Context Mode Gateway feature, keeps every word and sends the model your recent work plus an index of the rest. In a paired run on a 200K window, a long session cost 37.6% less than Claude Code direct, answered 9 of 10 questions about early details exactly against 4 of 10, and never compacted.

Key facts

What it is
A long-session layer in Context Mode Gateway: the model sees a bounded window of recent work and an index of older work, and the agent finds anything older in an archive with context-mode-search.
Who it is for
Developers who keep one Claude Code or Codex session open for hours and lose details when it compacts, and teams who pay for long sessions.
Clients
Recall works in Claude Code and Codex. On Codex, a later turn reading back an exact early detail from the archive is not proven yet. In every plan: see Plans and limits.
Measured result
A long Claude Code session on Claude Opus 5.5, 200K window, 3 pre-registered pairs on 2026-10-04: $3.04 through the gateway against $4.86 direct, 37.6% lower. Direct compacted 3 times; through the gateway the session never compacted. On the 1M window: 59.0% lower, with 10, 10, 10 and 8 of 10 memory answers against 10 of 10 direct. One live session (2026-10-02): 757,803 tokens per request before Recall, 150,638 where it held (JSON). Evaluation.
Limits
Cost measured in Claude Code on Claude Opus 5.5. Codex cost is not measured yet. The agent searches to see older work.

Summary

A long Claude Code session fills the model's context window. When it is close to full, Claude Code summarizes the older part and drops the details. Until then, every request sends the whole history again, so each request costs more than the one before.

Recall changes what the model gets, not what your client keeps. Your client still holds every message. The gateway sends the model your recent work, plus a short index of everything older. Before an older message leaves the model's view, the gateway archives it word for word. The agent can get it back by its id, or find it by searching. Nothing is summarized and nothing is thrown away.

What that did in a paired run of a long Claude Code session on a 200K window, against Claude Code direct (method):

And on the owner's own main session (method):

The paired runs were pre-registered: the rules were committed before the first pair. The owner's session and the probes are real sessions, not a benchmark. Every figure links to its evidence in Evaluation.

One long session, two endings

An example of the kind of session Recall is for. A developer works on billing all afternoon in one Claude Code session. Early on, they say: "Keep the refund window at 14 days. Do not touch the webhook retry code." The agent reads 40 files, runs the tests 30 times, and gets a screenshot of an error page. Four hours in, the session passes its window.

Four hours inClaude Code auto-compactClaude Code with Recall
Your rule "refund window 14 days"In the summary if the summary kept it, in other wordsIn the index word for word, and listed as a decision the agent must not reopen
"Do not touch the webhook retry code"Can get lost with the other early instructionsIn the index word for word
The exact failing line from test run 3Gone from the conversation. The agent runs the tests again to see it.One line in the index points to the archived run. The agent fetches it by id.
The error screenshotGoneArchived as an image. Its id in the index returns it.
The quota file it read at 14:10, edited at 15:30GoneThe old read comes back marked "changed at 15:30", with the id of the newer copy
Next requestStarts from the summaryCarries recent work and the index; the history in your client is untouched

The table is an illustration of the mechanism, not a measurement. The measured figures are in Evaluation.

The problem

Three things go wrong in a session that runs all day.

Memory tools help with the second point: they keep chosen facts. They do not shrink what the client sends, and a fact is not the exact line it came from.

How it works

Recall acts on the request between your client and the model. The client sends its full history, as it always does. The gateway decides what part of it the model gets.

When it acts

On most requests Recall changes nothing: the request goes out with the same start as the one before, so the provider reads it from its cache. It lays the request out again only at one of these moments.

MomentWhenWhy then
After a pauseLonger than the prompt cache lasts (5 minutes or 1 hour, as the client asks)The provider writes the whole prompt again anyway. A shorter prompt costs nothing extra to switch to, and is cheaper to write.
First timeThe first request of a thread that Recall has not laid out yetOne cache write, to start the bounded layout.
The windowThe request would pass your Context limitIt has to change to fit. This is the point where the client would otherwise compact.
The cost ruleWhile the cache is warm, the history keeps growingRecall adds up what the growing part costs to read on each request. When that sum reaches the cost of one cache write of the short layout, it lays out again. So the write is paid back before it happens.

It also lays out again when your client itself rewrites older messages, because the cache is written again in that case too.

What the model gets

Left: the client holds the system prompt, the tools and the full history. Right: the model gets the same system prompt and tools, then the memory block, then the index of older work, then the recent work unchanged. The older messages go to the archive, word for word. Your client holds system prompt and tools older messages hours of turns, files, output recent work The model gets system prompt and tools, unchanged memory: facts and decisions index of older work your words in full, one line per step recent work, unchanged archive
The older messages leave the model's view only. They stay in your client, and the archive holds an exact copy that the index points to.

From the top:

  1. The system prompt and tools, as your client sent them.
  2. The memory block: facts and decisions that are true now (see Decisions and Memory).
  3. The index of older work. One block of text, the same bytes on every request until the next layout, so it is read from the cache.
  4. Recent work: your client's own messages from the cut point to the end, unchanged. The cut point holds a fixed amount of recent work, set in Settings, and never moves on a request where nothing above happened.

The index

The index is built by rules, not by a model, so the same history always gives the same index. A short example:

Older part of this conversation, archived by context-mode. Nothing was deleted: every line names its exact archived copy.
Call context-mode-search with {"docid":"<id>"} for the full bytes, or with {"q":"..."} to search it.
2026-10-01 21:58 user: run the billing tests and fix what fails, keep the refund window at 14 days
  → Bash "npm test -- src/billing" (context-mode-msg-3f2a…)
  → Edit src/billing/quota.ts (context-mode-msg-9b1c…)
  ← assistant: Two tests failed in quota.ts. The refund window check used … (context-mode-msg-77de…)
2026-10-01 22:04 user: [img: context-mode-img-5c0e…] this is the error screen, fix it

Archived before it leaves

A message leaves the model's view only after the archive has confirmed it holds a copy. On the first layout of a very long thread, the archive may need more time than one request allows. Then that request goes out as it would have without Recall, the archive keeps filling in the background, and a later request makes the cut. Nothing leaves the model's view on an unconfirmed write.

If any check on the new layout fails, the gateway sends the request as it would have without Recall. It never sends a half-built layout.

Finding things: context-mode-search

The agent looks things up with the context-mode-search tool. The tool's own description tells the model to search before it asks you something you may have decided, before it says it does not remember, and before it runs a command again only to see an old result.

exact copy{"docid": "<id>"} returns the archived item as it was stored, paged with offset when long. An image id returns the image.
by words{"q": "refund window"} searches the archive and your facts.
scope"scope": this worktree by default; "project" for every worktree of the repository; "all" for everything on your account.
as of"as_of": "2026-10-01T22:00:00Z" asks what was known then: only items from before that time, and facts that were true at it.
kinds"kinds": ["msg", "tool", "img", "fact"] narrows the result.

Results come back with current facts first, then archived items ranked by how well they match, how close they are to this conversation, and how recent they are. Each item says where and when it happened.

"Changed since"

An old file read can be out of date. When the archive holds a read of a file and a later edit of the same file in the same worktree, the result says so: changed at 14:02, newer: <id>. A git checkout, pull, merge, rebase or reset marks the worktree's older file reads as "may have changed since". The tool tells the model to read the file on disk for its current content, and to prefer the newer item.

Pointers in your turn

When your new message names a file path or an error text that matches older archived work, the gateway adds one short line naming the archived item and its age. It adds the pointer, never the content, and at most 3 lines per turn.

Decisions

A decision you made once should not come back. Recall keeps decisions as facts, tied to the moment they came from.

In the console, decisions are listed under "Things you told your agent". Each one opens the moment it came from, and "Forget" closes it.

Scopes

LevelMeansShared by default
WorktreeOne topic of workYes: the default for search, decisions and pointers
Session, in one worktreeOne client processSeveral sessions in one worktree share the worktree's archive, facts and decisions. The current conversation ranks first.
ProjectThe repository, all its worktreesOnly when the search asks for "project". Each result names its worktree and branch.
AccountYouFacts about you, such as your name and how you like to work

So two worktrees of one project count as two topics: the decisions and pointers of one never show in the other unless you ask. Two sessions in one worktree work as one memory.

Recall follows the Memory scope: a search can only reach as far as the Memory page lets this worktree see, and tool output and files read stay in the worktree they came from.

Settings

Recent work in viewLean: newest 100K tokens. Balanced (default): newest 150K. Wide: newest 300K. Off: Recall does not act. Never more than half your Context limit. Older work is still searchable. Smaller is cheaper per request.
RetentionHow long the archive keeps older work: 30, 90 or 365 days, or forever. Default 90 days.
Team sharingShare a project's archive with your team. Off by default, and set per project.
Forget this projectDeletes the project's saved history and the facts learned from it. Your settings, other projects and usage stay. The confirm step names the project and how many items it deletes.

Balanced is the default because it keeps a full working stretch in view. A smaller window costs less per request.

What you see in the terminal

When Recall acts, the gateway adds one line to the reply, in the same banner as its other notes:

⚡ context-mode · picked up after a 2 h pause, nothing compacted
⚡ context-mode · recalled 2 items from 2 days ago (a decision, src/billing/quota.ts)

In the console, "Your agent looked back" lists each lookup ("the refund decision, found in billing, 2 h ago"), and a long session reads as one line: how many hours your agent remembers, and what each request carried against the full history.

What you save

The console shows a "saved by Recall" number on each request Recall touched, and the total under Saved. It compares two prices for the same request:

Saved by Recall is the second price minus the first. Only the model's first read of the request is compared; its reply and later calls cost the same either way.

Recall's own costs and failures count against it. A new layout costs one cache write. When Recall fails on a request, or lets go of a session, the full history is written to the cache again, so that request costs more than it would have without Recall. Those costs are subtracted, so the number can be below zero. It counts every main request from the first one Recall touched, so a loss after Recall stops is still Recall's. Subagents are not counted.

Try it in 5 minutes

A teammate changed a login message and two tests broke. A sub-agent ran the suite and sent back only the counts. Then you ask for the exact failing line, and nothing runs again.

Setup

npx @context-mode/cli      # sign in and connect Claude Code, then restart Claude Code
git clone https://github.com/expressjs/express.git && cd express && npm install
# play the teammate: a copy change to the login error
sed -i.bak "s/'Authentication failed, please check your '/'Login failed, please check your '/" examples/auth/index.js
git commit -qam 'copy: friendlier login error'
claude

Prompts

One session, one prompt at a time:

Use a subagent to run the full test suite (npm test). I only need the pass and fail counts back from it.
Which assertion failed in the auth tests? Quote the exact error line. Don't run the tests again.

What you will see

# prompt 1: counts only. The sub-agent read the log; your session did not.

# prompt 2: no test run, the exact lines
test/acceptance/auth.js:36:10
to match /Authentication failed/
      at Test.<anonymous> (test/acceptance/auth.js:36:10)

test/acceptance/auth.js:51:14
to match /Authentication failed/
      at Test.<anonymous> (test/acceptance/auth.js:51:14)

Both assertions expect the response body to contain /Authentication failed/, but the actual response contains "Login failed, please check your username and password..." instead.

⏺ context-mode · recalled 10 items from 1 minute ago (a fact, Bash)

Two worktrees of one repo

Now open a second worktree of the same repo, for example with claude --worktree router-cache, and ask:

Paste the exact error from the auth test run earlier today. Don't run the tests.

The agent says it has no record of that run. Tool output, files read and logs stay in the worktree where they happened. Another branch may have different code, so an old log from there could mislead the agent.

The next session

Close the session, start a new one in the same folder, and ask:

Paste the exact error from the auth test run before lunch, I need it for the bug ticket. Don't run the tests again.

On the first message of a new session, the gateway searches the archive itself and puts the matching lines of the earlier run in front of the agent. The agent quotes the two failing assertions from the earlier run without running anything: 5 of 5 runs.

We ran this five times on 2026-10-04 in real Claude Code sessions (claude -p 2.1.283, Claude Haiku 4.5) through the gateway, each on a new test account. Prompt 2 quoted the exact line in 5 of 5 runs, and so did the next session. Recall is on for every account by default; you turn it off on the Memory page.

Why it matters. The exact line you need is still there hours later, even when a sub-agent read it. You do not run a slow suite again just to see one line.

Recall and Memory together

Recall and Memory are two halves of one memory.

MemoryThe current truth: short facts with two clocks, when each was true and when the gateway learned it.
RecallThe evidence: every message and tool call as it happened, with when, where and who.

Each fact links to the moment it came from. When a fact changes, the old one is closed, and the archived moment that set it stays. So the agent can answer "what is true now" from Memory, and "where did that come from" or "what did the file say then" from Recall.

Evaluation

Two pre-registered paired runs of a long Claude Code session, an offline replay, live checks and the owner's own session. Every number links to the file it comes from.

Long session on a 200K window: 37.6% lower cost, 9 of 10 exact answers

Long session on the 1M window: 59.0% lower cost

What it means for you. A long session costs less than it does direct, and an early detail comes back from the archive, not from a summary.

Quality replay, no model

Live checks

Read speed

The owner's main session

Recall probe: deleted tool output

Decisions check

Verify it yourself

Every evidence file holds the token counts of each call and the list price we used, so you can recompute every dollar yourself. A checking script is available on request.

Compared with other approaches

From each vendor's own pages, read on 2026-10-02. "Not stated" means the page we read does not say.

ApproachKeeps every detail recallableExact copyKnows what changedCost on a long sessionNo change to how you work
Context Mode RecallYes. Every older message and tool call is archived, word for word except masked secrets and e-mail addresses.Yes, by id. Images come back as images.Yes. A file read later edited says "changed at"; "as of" asks what was known then.Each request carries your recent work and the index, not the whole history.Yes, after one setup command.
Claude Code auto-compactNo. Older history becomes a summary; early instructions can get lost.NoNoGrows with every turn until the window, then drops to the summary.Yes, it is built in.
Memory tools: mem0, supermemory, claude-memWhat the tool chose to keep: facts, notes or saved conversationsNot statedNot statedThe client still sends its full history and compacts at its window. The tool adds its own tokens.A plugin or hooks on each machine.

Where Context Mode differs. Recall does not choose what to keep. It keeps everything, points to it from the index, and lets the agent fetch the exact item. It also shortens what each request carries, which a memory tool does not do.

Where others are stronger.

Privacy and retention

Limits

FAQ for CTOs

Does Recall compact my conversation?

No. Your client keeps its full history, and the gateway never summarizes it. The model gets your recent work and an index; everything older is in the archive, word for word.

Is anything deleted?

Not while it is in your retention period (default 90 days). Older messages leave the model's view, not the archive. Only masked secrets and e-mail addresses cannot be read back.

Does it work with Codex?

Yes. Recall works in Claude Code and Codex. On Codex, a later turn reading back an exact early detail from the archive is not proven yet.

Does it cost less?

Yes, on a long session. In the paired runs, a long session through the gateway with Recall on cost 37.6% less than Claude Code direct on a 200K window, and 59.0% less on the 1M window. Each request carries fewer tokens: 150,638 instead of 757,803 on the owner's session. A new layout costs one cache write, and Recall makes one only after a pause, the first time, at the window, or when the cost rule says it has paid for itself.

Can I turn it off?

Yes. Set "Recent work in view" to Off in Settings. Your archive stays until its retention ends, and "Forget this project" deletes it now.

Can I get an old screenshot back?

Yes. A pasted image is archived as an image, and its id in the index returns it.

Compared with other tools: see the landscape.