Wrong skills, noisy memory, false resumes: fixing the small choices behind every turn

Engineering6 min read

Your coding agent loads the wrong skill, gets facts it does not need, or replays yesterday's session after you type "hi". Each of these is a small choice Context Mode makes for your agent on every turn. A text model used to make them, and it was often wrong and sometimes very slow. We moved four of these choices to Cloudflare's Clef decision models, measured each move, and shipped three. In an offline test, the right skill fired 9 or 10 times in 10 instead of 4, and memory sent 41 tokens per opening instead of 262.

The problem: small wrong choices you pay for all day

If you run a coding agent all day, you have seen these:

Context Mode sits between your coding agent and the model, so these choices are ours to get right. Here is what we changed.

The fix: a model that picks, not one that writes

Before, a text model (llama-3.3-70b-instruct-fp8-fast) got a prompt and had to write back an answer, such as one skill id or NONE. Clef works differently. You send a typed question, either "pick one of these options" or "yes or no", and you get back a probability for every option.

A number per option is what makes the fix work. We can set a floor: if no option is likely enough, nothing fires, and we can log why.

How we measured

1. The right skill: 4 of 10 became 9 or 10 of 10

On the 10 disputed turns where the rule decides the answer:

JudgeRight
70B text model (old)4 of 10
Clef 27B, same candidates, fire at p ≥ 0.59 of 10
Clef 27B, whole catalog, fire at p ≥ 0.510 of 10

The bigger change for you is speed at the slow end:

70B text model, p954,728 ms
Clef 27B, p95612 ms
Skill choice time inside the Worker, 20 calls per arm. p50 went from 397 ms to 281 ms.

The floor matters. Clef's top answer fires on one-word test prompts at p 0.3 to 0.4. Without the 0.5 floor, it fired on 21 of 121 turns and was right less often than the old model. With the floor, those fires stop.

Asking over the whole catalog also fixes search misses. The old route first searched for at most 8 candidate skills. On 41 made-up labeled cases, 3 of its 4 misses came from that search finding nothing. The whole-catalog question has no search step and got all 3: 39 of 41 against 37 of 41. That set is small, so it does not prove the two apart on its own.

What shipped: Clef 27B over the whole catalog, fire at p ≥ 0.5, with a hard stop at 2,000 ms.

Limit: these numbers use an 11-skill test catalog, because the owner's tenants have no skills.

2. Memory: fewer facts, and more of them useful

At the start of a conversation, Context Mode sends a short block of facts it remembers about you. Before, it sent about six rows every time, mostly your newest facts, whatever the turn was about.

Now clef-flash reads a pool of 20 rows and asks, for each one, "is this relevant to this turn?" Rows at p ≥ 0.25 ship, best first, at most six.

Before262
clef-flash filter41
Memory tokens added per conversation opening, on 52 openings (offline test).
On 52 conversation openingsBeforeclef-flash filter
Rows sent that were relevant (precision)4.2%45.3%
Relevant rows that were found (recall)27.7%61.7%
Rows sent5.961.23
Memory tokens added26241

A 70B judge labeled relevance, twice, in two row orders. On the stricter labels (rows both passes agree on), precision went from 1.6% to 14.1% and recall from 27.8% to 50.0%. Same verdict.

This one adds a step on our side: one extra clef-flash call per opening, with p50 626 ms in the Worker and a hard stop at 1,200 ms. On a timeout or error, the old six rows ship. Your side gets about 200 to 220 fewer input tokens on that opening request. Later turns pay nothing.

3. Resume: every "yes" was wrong

When your first message looks like "continue where we left off", Context Mode replays your last session. A similarity score against 10 sample phrases used to decide this.

What shipped: when the old score says yes, clef-flash gets one yes/no question, with a hard stop at 300 ms. On a timeout, the old answer stands. The extra call ran on 2 of 141 human turns.

Limit: the test set held no real resume request, so we have not measured how many real ones it catches. We tested Turkish and English only.

4. The one we did not ship

After each turn, a 70B model reads it and pulls out facts to remember. We tried a Clef yes/no question in front of it: "does this turn state a lasting fact?" The best version sent 68% of typed turns to the 70B and lost 1 fact turn of 69. It saved very little.

We left it off. The working threshold sat just above the lowest score Clef gives on this question, so a small model update could change what it catches. And to keep every fact turn, the gate had to send 90% to 99% of turns, so it saved almost nothing. It stays off by default.

5. The biggest saving came from a plain rule

A turn with no words you typed, such as a tool result or a subagent's last turn, can never produce a stored fact. We now skip the 70B call on those turns, with no model involved, and a test runs the real filter to prove nothing is lost.

Correctness and speed decided these moves, not price.

Where a decision model helps, and where it does not

We chose the 0.5 skill floor and the 0.25 memory cut after looking at the data, so each change has its own off switch.

Not measured yet

How to try it

Run one command:

npx @context-mode/cli

It sets up Claude Code or Codex to send requests through Context Mode. Start free with 1,000 requests, every feature, no card. Your own Anthropic or OpenAI account pays for the model.

More on the request path these choices run in: Slipstream: a shorter path from your agent to the model.