# All-day agent sessions without compaction: how Recall keeps every word

October 9, 2026 Recall 6 min read

Keep one coding agent session open all day and it ends one of two ways: the client compacts and the agent forgets your morning, or you pay to resend the whole history on every request. Recall, part of Context Mode, gives you a third way. It moves older work into an exact archive, not a summary, and the agent reads it back word for word. In a pre-registered A/B, the session through Context Mode never compacted, answered 9 of 10 early-detail questions against 4 of 10, and cost 37.6% less.

## The problem: your agent forgets the morning

You know the moment. The context window fills, the client compacts, and the agent now works from a summary of the last few hours.

A summary is short on purpose, so it drops things:

- the exact failing line from the third test run,
- the reason you chose option B over option A,
- a file the agent read at 10:00 that is no longer on disk,
- an early instruction such as "do not touch the webhook retry code".

Claude Code's own docs say early instructions can get lost when it compacts. Then the agent asks you the same question twice, or runs a command again just to see an old result.

## The second problem: you pay for the whole history, every time

Before the window fills, each request sends the full history again. In one of our own long sessions, each request carried about 700K to 740K tokens.

- While the prompt cache was warm, a request cost $0.16 to $0.18.
- After a pause of more than one hour, a request cost $5.86 to $5.90, because the provider writes the whole prompt to its cache again.

So a long session has two bad endings: lose details, or pay for every token after every long pause.

## How Recall fixes it

Context Mode is the control plane between your coding agent and the model. Recall runs there, in the request path. Your client still sends and keeps its full history. Context Mode decides which part the model sees.

When Recall acts, the model gets this, from the top:

1. The system prompt and tools, as your client sent them.
2. A memory block: facts and decisions that are true now.
3. An index of older work. Its bytes stay the same on every request until the next layout, so the provider reads it from cache.
4. Recent work, unchanged. The default setting, "Balanced", keeps the newest 150K tokens in view.

The index is built by fixed rules, not by a model, so the same history always gives the same index. Every word you typed stays in it, verbatim. Each tool call becomes one line with its main input and an id:

```
2026-10-01 21:58 user: run the billing tests and fix what fails, keep the refund window at 14 days
→ Bash "npm test -- src/billing" (context-mode-msg-3f2a…)
→ Edit src/billing/quota.ts (context-mode-msg-9b1c…)
← assistant: Two tests broke in quota.ts. The refund window check used … (context-mode-msg-77de…)
```

## Exact text back, on demand

The agent gets older work back with one tool, `context-mode-search`:

- `{"docid": "context-mode-msg-3f2a…"}` returns the archived item byte for byte, except secrets and e-mail addresses, which are masked before storage.
- `{"q": "refund window"}` searches the archive and your facts.

If the archive holds a read of a file and a later edit of that file, the result says `changed at 14:02, newer: <id>`. The agent does not trust a stale copy.

Two safety rules:

- A message leaves the model's view only after the archive confirms it holds an exact copy.
- If a check on the new layout fails, Context Mode sends the request as it would have without Recall. It never sends a half-built layout.

## When it acts, and when it stays out of the way

On most requests Recall changes nothing, so your prompt cache keeps working. It builds a new layout only at three moments:

- **After a long pause.** The cache has expired and the provider will write the whole prompt again anyway, so a shorter prompt costs nothing extra.
- **At the window.** When the request would pass your Context limit, the point where the client would otherwise compact.
- **By a cost rule.** While the cache is warm, when the extra reads of the growing history reach the cost of one cache write of the short layout.

## The proof: a pre-registered A/B

We fixed the rules on 2026-10-04 before the first pair ran. Both arms used Claude Code 2.1.283 on Claude Opus 5.5 with a real 200,000-token window. Each ran the same 30-step coding day: plant four facts, read reference files, print three reports that are then deleted, fix three tests, add a CSV export, then answer ten questions about early details.

One arm talked to Anthropic directly. The other went through Context Mode, with the Context limit set to 100,000 so Recall would act. Dollars are Anthropic list price applied to the token counts Anthropic returned.

**Claude Code direct** *4 of 10* **Through Context Mode** *9 of 10*

Exact answers to ten questions about early details, mean of 3 pairs. Through Context Mode, pairs ranged from 7 to 10.

| Mean of 3 pairs | Claude Code direct | Through Context Mode |
| --- | --- | --- |
| Cost per run | $4.86 | $3.04 (37.6% less) |
| Exact answers, of 10 | 4 | 9 (7 to 10) |
| Compactions | 3 | 0 |
| Task checks, of 8 | 8 | 8 |
| Tokens sent per call, mean | 102,186 | 48,743 |

Direct Claude Code compacted three times. It lost every deleted report and every constant from the files it read. Through Context Mode, the deleted reports came back from the archive in every pair. Both arms passed all 8 task checks.

## Where the saving really comes from

Recall's job is to keep the details, not to cut the bill on a short run. On this run, Recall's own line in our ledger was negative: the new layout's cache write cost about $1.03 more than it saved. Most of the 37.6% came from the rest of Context Mode, such as smaller tool outputs, and from never compacting. In an earlier run where Recall never acted, the Context Mode arm cost $2.30 and answered 10 of 10.

It is also slower per request. Time to first token, p50, was 4.6 s through Context Mode against 2.2 s direct. The 30 steps took 393 s against 329 s.

More results:

- **1M window** (separate run, 4 pairs, 2026-10-03): the Context Mode arm cost 59.0% less. It answered 10, 10, 10 and 8 of 10 exactly. Direct answered 10 of 10 every time, because at 1M it never had to compact.
- **One real session, not a benchmark:** before Recall acted, each request sent a mean of 757,803 tokens. Where Recall held, the mean was 150,638 tokens, with no upstream errors and no compaction in that window.

## What is not measured yet

- Recall at the default Context limit, on a workload long enough to cross it by itself. Our 200K run lowered the limit to get there.
- How much of the saving comes from which feature. The 200K run had no arm with Context Mode but without Recall.
- Codex. Recall works in Codex, but reading back an exact early detail on a later Codex turn is not proven yet, and Recall's effect on Codex cost is not measured yet.
- Sample size: 3 pairs on 200K, 4 pairs on 1M, one model.

## One bug it taught us

On 2026-10-08, from 08:33 to 11:04 UTC, turns that used Recall waited as long as 15 s. Between 08:33 and 10:40, all 201 Recall reads returned an error. The cause was one line in a wrapper around a Durable Object stub:

```
// before
(a) => v.apply(t, a)
// after
(a) => Reflect.apply(v, t, a)
```

A Cloudflare RPC stub answers every property access with another stub. So `.apply` there is a remote call to a method named "apply", which does not exist. `.bind` and `.call` fail the same way. We fixed all 15 call sites in 11 files and added a regression test. After the fix, 47 of 47 Recall reads succeeded, p50 9 ms. More in [Slipstream](https://context-mode.com/blog/slipstream).

## How to try it

```
npx @context-mode/cli
```

- It signs you in and writes one config change per agent, Claude Code or Codex. Node 20.12 or later.
- Recall is on in every plan, and you can start free with no card. See [Plans and limits](https://context-mode.com/docs/plans-and-limits).
- Your own Anthropic or OpenAI account pays the model.
- To undo it: `npx @context-mode/cli disconnect`.

Start with the [Quick start](https://context-mode.com/docs/quick-start). Read how Recall works on the [Recall page](https://context-mode.com/recall) and in the [Recall docs](https://context-mode.com/docs/recall).

[Start free](https://context-mode.com/docs/quick-start)[All posts](https://context-mode.com/blog)
