Agent memory that knows when a fact changed

Memory6 min read

A coding agent starts each session with what its client loads, and nothing more. Memory in Context Mode takes short facts from your turns, stores each one with two clocks (when it was true, and when we learned it), and adds only the facts that matter to the first request of a new conversation, in Claude Code and Codex. In an offline eval, memory tokens per new conversation fell from 262 to 41, and the share of relevant facts rose from 4.2% to 45.3%.

The problem: you say it twice, or the agent says it wrong

You tell your agent the deploy target on Monday. On Tuesday, in a new session, it asks again.

The fix most tools offer is a memory store that only adds rows. That makes a second problem. You moved hosts on Monday afternoon, so now the store holds two "current" answers, and the agent may pick the old one.

A useful memory has to do three things well:

Here is how Context Mode does each one, and what we measured. Some samples are small; we say which.

The short version

Facts that change: two clocks, not two answers

Every fact is a short row, such as production deployed_to "Render, Oregon region". Each row carries two time ranges:

ClockAnswers
Valid timeWhen was this true? An empty end means "true now".
Recorded timeWhen did the store learn it, and when did it close it?

When a new value arrives for the same slot (the same subject and predicate), Context Mode does not overwrite. It closes the old row at the moment the new one starts and adds the new row. The two windows meet with no gap and no overlap.

The hard part was knowing which facts hold one value. Our first version used a hand-written list of 17 predicates, and everything else was added on top. On our own graph, that meant the staging database's region was both Reykjavik and Frankfurt, both marked valid. The list could not keep up: the extractor coined 81 predicates across 108 facts.

So we turned the default around. A fact holds one value unless it is declared to hold many, or it is an event or a two-way relation. Every close is recorded and can be undone.

In a store eval of 24 facts that changed 38 times, using the gateway's own write and read code:

CheckContext Mode (closed rows)Simple store
Memory blocks holding an old value0 of 2423 of 24 (append-only)
"What was true on this date" questions answered38 of 380 of 38 (a store that overwrites)

This is a controlled test of the mechanism on one scenario, not a test of recall in real use.

Facts in the right place: whose past is this?

The bug that made us build this: in two worktrees of the same repo, the agent answered "what is our release codename, and what file did you read first?" from the other worktree's past.

Context Mode now gives every stored item a place:

The filter runs before anything reaches the model, so the model never sees an item from the wrong place. A live check with throwaway accounts, where worktree A states a decision and reads a file:

CaseClaude Code, run 1Claude Code, run 2
Worktree B, new session, knows the decisionyesyes
Worktree B knows A's filenono
Another repo knows eithernono

One Codex run gave the same split. Known gaps: a worktree outside the known folder markers counts as its own project, and a client that replaces its own system prompt sends no folder, so its read covers the whole account.

Fewer facts, better facts

Until our v3 release, each new conversation got the six top-ranked facts. On 161 recorded turns, 913 of the 998 ranking tags on those rows meant "the user's newest facts". So the block held your latest facts, whatever the turn was about. 4% of them were relevant.

Now Context Mode reads a pool of 20 facts and asks Cloudflare's clef-flash model one yes/no question per fact: does this fit the turn? It keeps at most six.

Before: six newest-ranked facts262
Now: clef-flash keeps the ones that fit41
Memory tokens per new conversation, on 52 conversation openings (offline eval, one tenant).
On 52 conversation openingsBeforeNow
Facts sent5.961.23
Share of sent facts that were relevant4.2%45.3%
Share of relevant facts found27.7%61.7%
Memory tokens per opening26241

The check takes 626 ms at p50. If it runs past 1.2 s, it stops and the old six facts go out, so a slow check never blocks your turn.

Read these numbers as direction, not as a final score. They come from one tenant's recorded turns, and the model that judged "relevant" agrees with itself only so-so across two passes (Cohen's kappa 0.37). On the strict label, where both passes agree, relevance went from 1.6% to 14.1%. Both labels point the same way.

Why the fact writer is still a 70B model

After a turn ends, off your request path, llama-3.3-70b-instruct-fp8-fast on Cloudflare Workers AI reads the turn and writes the facts. It is the most expensive model step we run, so we tried hard to replace it.

We set the bar before the test: a cheaper model wins only if it names the same facts as the 70B on at least 90% of 300 recorded turns.

ModelAgreement, 300 turnsSame facts on the 54 turns that held one
70B against itself (noise floor)97.3%87.0%
glm-4.7-flash83.0%16.7%
qwen3-30b-a3b82.3%16.7%
llama-4-scout-17b82.0%13.0%
gemma-4-26b-a4b80.0%3.6%
granite-4.0-h-micro72.3%3.6%

None passed. Scores near 80% look close, but most of that is turns where both models found nothing. On turns that hold a fact, the cheaper models named the same fact 13% to 17% of the time. They do see that a fact is there (glm-4.7-flash agrees with the 70B on that 95.7% of the time), but they word it differently or pick another one. A memory that closes old values by slot needs the same wording every time.

Other ways to cut cost fell short too:

What we skip instead. Skipping short turns does not work: the same one-word reply, yes, gave 0 lasting facts in one turn and 5 in another, because yes confirmed facts in the agent's answer. What we do skip is turns with no words a person typed, such as a tool result or a subagent's last turn. Those can never keep a fact, so the 70B call is skipped, and a test proves nothing is lost.

How much this saves in production is not measured yet. In the recorded eval set, 246 of 379 turn pairs had no typed words and 0 of them kept a fact, but that set includes tool results, so it overstates the real share.

What is not measured yet

How to try it

npx @context-mode/cli

It signs you in and connects Claude Code or Codex. OpenCode and Pi are coming soon. Memory is in every plan, and your own Anthropic or OpenAI account pays the model. Plan limits are on Plans and limits.