# Agent memory that knows when a fact changed

October 9, 2026 Memory 6 min read

A coding agent starts each session with what its client loads, and nothing more. Memory in Context Mode takes short facts from your turns, stores each one with two clocks (when it was true, and when we learned it), and adds only the facts that matter to the first request of a new conversation, in Claude Code and Codex. In an offline eval, memory tokens per new conversation fell from 262 to 41, and the share of relevant facts rose from 4.2% to 45.3%.

## The problem: you say it twice, or the agent says it wrong

You tell your agent the deploy target on Monday. On Tuesday, in a new session, it asks again.

The fix most tools offer is a memory store that only adds rows. That makes a second problem. You moved hosts on Monday afternoon, so now the store holds two "current" answers, and the agent may pick the old one.

A useful memory has to do three things well:

- **Remember** the facts you state, without you writing them down.
- **Notice when a fact stopped being true**, and stop sending the old value.
- **Send back only what fits the turn**, so memory does not fill the context with noise.

Here is how Context Mode does each one, and what we measured. Some samples are small; we say which.

## The short version

- When a fact changes, the old row is closed, not kept as a second "current" answer. In a store eval, 0 of 24 memory blocks held an old value, against 23 of 24 for an append-only store.
- Facts stay in the right place. A decision from one worktree reaches the other worktree of the same repo; file reads and other repos stay out.
- A new conversation gets fewer, better facts: 1.23 rows instead of 5.96, and 41 memory tokens instead of 262 (offline eval, one tenant).
- We tried five cheaper models to write the facts. None was good enough, so the extractor is still a 70B model. We skip it on turns where it cannot keep anything.

## Facts that change: two clocks, not two answers

Every fact is a short row, such as `production deployed_to "Render, Oregon region"`. Each row carries two time ranges:

| Clock | Answers |
| --- | --- |
| Valid time | When was this true? An empty end means "true now". |
| Recorded time | When did the store learn it, and when did it close it? |

When a new value arrives for the same slot (the same subject and predicate), Context Mode does not overwrite. It closes the old row at the moment the new one starts and adds the new row. The two windows meet with no gap and no overlap.

- A value that arrives late, with an older date, is stored already closed, as history.
- Nothing is deleted. A slot closed by mistake can be reopened with one call.
- You can see the facts, the closed rows and both clocks on the Memory screen in the console.

**The hard part** was knowing which facts hold one value. Our first version used a hand-written list of 17 predicates, and everything else was added on top. On our own graph, that meant the staging database's region was both Reykjavik and Frankfurt, both marked valid. The list could not keep up: the extractor coined 81 predicates across 108 facts.

So we turned the default around. A fact holds one value unless it is declared to hold many, or it is an event or a two-way relation. Every close is recorded and can be undone.

In a store eval of 24 facts that changed 38 times, using the gateway's own write and read code:

| Check | Context Mode (closed rows) | Simple store |
| --- | --- | --- |
| Memory blocks holding an old value | 0 of 24 | 23 of 24 (append-only) |
| "What was true on this date" questions answered | 38 of 38 | 0 of 38 (a store that overwrites) |

This is a controlled test of the mechanism on one scenario, not a test of recall in real use.

## Facts in the right place: whose past is this?

The bug that made us build this: in two worktrees of the same repo, the agent answered "what is our release codename, and what file did you read first?" from the *other* worktree's past.

Context Mode now gives every stored item a place:

- **Text you typed** in the main thread is shared across the project (the repository).
- **Tool output and file reads** stay in their worktree.
- **Anything that looks like a secret** stays in its thread.

The filter runs before anything reaches the model, so the model never sees an item from the wrong place. A live check with throwaway accounts, where worktree A states a decision and reads a file:

| Case | Claude Code, run 1 | Claude Code, run 2 |
| --- | --- | --- |
| Worktree B, new session, knows the decision | yes | yes |
| Worktree B knows A's file | no | no |
| Another repo knows either | no | no |

One Codex run gave the same split. Known gaps: a worktree outside the known folder markers counts as its own project, and a client that replaces its own system prompt sends no folder, so its read covers the whole account.

## Fewer facts, better facts

Until our v3 release, each new conversation got the six top-ranked facts. On 161 recorded turns, 913 of the 998 ranking tags on those rows meant "the user's newest facts". So the block held your latest facts, whatever the turn was about. 4% of them were relevant.

Now Context Mode reads a pool of 20 facts and asks Cloudflare's clef-flash model one yes/no question per fact: does this fit the turn? It keeps at most six.

**Before: six newest-ranked facts** *262* **Now: clef-flash keeps the ones that fit** *41*

Memory tokens per new conversation, on 52 conversation openings (offline eval, one tenant).

| On 52 conversation openings | Before | Now |
| --- | --- | --- |
| Facts sent | 5.96 | 1.23 |
| Share of sent facts that were relevant | 4.2% | 45.3% |
| Share of relevant facts found | 27.7% | 61.7% |
| Memory tokens per opening | 262 | 41 |

The check takes 626 ms at p50. If it runs past 1.2 s, it stops and the old six facts go out, so a slow check never blocks your turn.

Read these numbers as direction, not as a final score. They come from one tenant's recorded turns, and the model that judged "relevant" agrees with itself only so-so across two passes (Cohen's kappa 0.37). On the strict label, where both passes agree, relevance went from 1.6% to 14.1%. Both labels point the same way.

## Why the fact writer is still a 70B model

After a turn ends, off your request path, `llama-3.3-70b-instruct-fp8-fast` on Cloudflare Workers AI reads the turn and writes the facts. It is the most expensive model step we run, so we tried hard to replace it.

We set the bar before the test: a cheaper model wins only if it names the same facts as the 70B on at least 90% of 300 recorded turns.

| Model | Agreement, 300 turns | Same facts on the 54 turns that held one |
| --- | --- | --- |
| 70B against itself (noise floor) | 97.3% | 87.0% |
| glm-4.7-flash | 83.0% | 16.7% |
| qwen3-30b-a3b | 82.3% | 16.7% |
| llama-4-scout-17b | 82.0% | 13.0% |
| gemma-4-26b-a4b | 80.0% | 3.6% |
| granite-4.0-h-micro | 72.3% | 3.6% |

None passed. Scores near 80% look close, but most of that is turns where both models found nothing. On turns that hold a fact, the cheaper models named the same fact 13% to 17% of the time. They do see that a fact is there (glm-4.7-flash agrees with the 70B on that 95.7% of the time), but they word it differently or pick another one. A memory that closes old values by slot needs the same wording every time.

Other ways to cut cost fell short too:

- A small 8B model recovered 5 of 103 facts and did not hold the output format.
- Sending five turns per call cut cost 63% but lost 42.8% of lasting facts.
- Prompt caching is not offered on this model.

**What we skip instead.** Skipping short turns does not work: the same one-word reply, `yes`, gave 0 lasting facts in one turn and 5 in another, because `yes` confirmed facts in the agent's answer. What we do skip is turns with no words a person typed, such as a tool result or a subagent's last turn. Those can never keep a fact, so the 70B call is skipped, and a test proves nothing is lost.

How much this saves in production is not measured yet. In the recorded eval set, 246 of 379 turn pairs had no typed words and 0 of them kept a fact, but that set includes tool results, so it overstates the real share.

## What is not measured yet

- Memory recalled on a later day, in real use, and in Codex.
- How many production extraction calls the skip for untyped turns removes.
- Cost. In a small live eval with Claude Opus 5.5, 8 probes cost $0.184 with memory on against $0.069 with memory off, as the client reported. We have not found the cause. In earlier runs on Claude Haiku 4.5, before the clef-flash filter, memory added 534 to 548 uncached input tokens to a new session's first request.

## How to try it

```
npx @context-mode/cli
```

It signs you in and connects Claude Code or Codex. OpenCode and Pi are coming soon. Memory is in every plan, and your own Anthropic or OpenAI account pays the model. Plan limits are on [Plans and limits](https://context-mode.com/docs/plans-and-limits).

- [Quick start](https://context-mode.com/docs/quick-start): install and connect your agent.
- [Memory](https://context-mode.com/memory): what it stores and how you see it.
- [Memory docs](https://context-mode.com/docs/memory): the two clocks, scope and the live eval in detail.

[Start free](https://context-mode.com/docs/quick-start)[All posts](https://context-mode.com/blog)
