All-day agent sessions without compaction: how Recall keeps every word

Recall6 min read

Keep one coding agent session open all day and it ends one of two ways: the client compacts and the agent forgets your morning, or you pay to resend the whole history on every request. Recall, part of Context Mode, gives you a third way. It moves older work into an exact archive, not a summary, and the agent reads it back word for word. In a pre-registered A/B, the session through Context Mode never compacted, answered 9 of 10 early-detail questions against 4 of 10, and cost 37.6% less.

The problem: your agent forgets the morning

You know the moment. The context window fills, the client compacts, and the agent now works from a summary of the last few hours.

A summary is short on purpose, so it drops things:

Claude Code's own docs say early instructions can get lost when it compacts. Then the agent asks you the same question twice, or runs a command again just to see an old result.

The second problem: you pay for the whole history, every time

Before the window fills, each request sends the full history again. In one of our own long sessions, each request carried about 700K to 740K tokens.

So a long session has two bad endings: lose details, or pay for every token after every long pause.

How Recall fixes it

Context Mode is the control plane between your coding agent and the model. Recall runs there, in the request path. Your client still sends and keeps its full history. Context Mode decides which part the model sees.

When Recall acts, the model gets this, from the top:

  1. The system prompt and tools, as your client sent them.
  2. A memory block: facts and decisions that are true now.
  3. An index of older work. Its bytes stay the same on every request until the next layout, so the provider reads it from cache.
  4. Recent work, unchanged. The default setting, "Balanced", keeps the newest 150K tokens in view.

The index is built by fixed rules, not by a model, so the same history always gives the same index. Every word you typed stays in it, verbatim. Each tool call becomes one line with its main input and an id:

2026-10-01 21:58 user: run the billing tests and fix what fails, keep the refund window at 14 days
→ Bash "npm test -- src/billing" (context-mode-msg-3f2a…)
→ Edit src/billing/quota.ts (context-mode-msg-9b1c…)
← assistant: Two tests broke in quota.ts. The refund window check used … (context-mode-msg-77de…)

Exact text back, on demand

The agent gets older work back with one tool, context-mode-search:

If the archive holds a read of a file and a later edit of that file, the result says changed at 14:02, newer: <id>. The agent does not trust a stale copy.

Two safety rules:

When it acts, and when it stays out of the way

On most requests Recall changes nothing, so your prompt cache keeps working. It builds a new layout only at three moments:

The proof: a pre-registered A/B

We fixed the rules on 2026-10-04 before the first pair ran. Both arms used Claude Code 2.1.283 on Claude Opus 5.5 with a real 200,000-token window. Each ran the same 30-step coding day: plant four facts, read reference files, print three reports that are then deleted, fix three tests, add a CSV export, then answer ten questions about early details.

One arm talked to Anthropic directly. The other went through Context Mode, with the Context limit set to 100,000 so Recall would act. Dollars are Anthropic list price applied to the token counts Anthropic returned.

Claude Code direct4 of 10
Through Context Mode9 of 10
Exact answers to ten questions about early details, mean of 3 pairs. Through Context Mode, pairs ranged from 7 to 10.
Mean of 3 pairsClaude Code directThrough Context Mode
Cost per run$4.86$3.04 (37.6% less)
Exact answers, of 1049 (7 to 10)
Compactions30
Task checks, of 888
Tokens sent per call, mean102,18648,743

Direct Claude Code compacted three times. It lost every deleted report and every constant from the files it read. Through Context Mode, the deleted reports came back from the archive in every pair. Both arms passed all 8 task checks.

Where the saving really comes from

Recall's job is to keep the details, not to cut the bill on a short run. On this run, Recall's own line in our ledger was negative: the new layout's cache write cost about $1.03 more than it saved. Most of the 37.6% came from the rest of Context Mode, such as smaller tool outputs, and from never compacting. In an earlier run where Recall never acted, the Context Mode arm cost $2.30 and answered 10 of 10.

It is also slower per request. Time to first token, p50, was 4.6 s through Context Mode against 2.2 s direct. The 30 steps took 393 s against 329 s.

More results:

What is not measured yet

One bug it taught us

On 2026-10-08, from 08:33 to 11:04 UTC, turns that used Recall waited as long as 15 s. Between 08:33 and 10:40, all 201 Recall reads returned an error. The cause was one line in a wrapper around a Durable Object stub:

// before
(a) => v.apply(t, a)
// after
(a) => Reflect.apply(v, t, a)

A Cloudflare RPC stub answers every property access with another stub. So .apply there is a remote call to a method named "apply", which does not exist. .bind and .call fail the same way. We fixed all 15 call sites in 11 files and added a regression test. After the fix, 47 of 47 Recall reads succeeded, p50 9 ms. More in Slipstream.

How to try it

npx @context-mode/cli

Start with the Quick start. Read how Recall works on the Recall page and in the Recall docs.