Wrong skills, noisy memory, false resumes: fixing the small choices behind every turn
Your coding agent loads the wrong skill, gets facts it does not need, or replays yesterday's session after you type "hi". Each of these is a small choice Context Mode makes for your agent on every turn. A text model used to make them, and it was often wrong and sometimes very slow. We moved four of these choices to Cloudflare's Clef decision models, measured each move, and shipped three. In an offline test, the right skill fired 9 or 10 times in 10 instead of 4, and memory sent 41 tokens per opening instead of 262.
The problem: small wrong choices you pay for all day
If you run a coding agent all day, you have seen these:
- The wrong skill loads, or the right one does not. The agent follows instructions that do not fit your task.
- Memory sends facts that do not matter. You pay tokens for them on every new conversation, and they can pull the model off course.
- A greeting replays your last session. You type "hi" or "run the tests", and the agent starts on yesterday's work.
- A slow choice holds up your turn, or times out and silently becomes "no skill".
Context Mode sits between your coding agent and the model, so these choices are ours to get right. Here is what we changed.
The fix: a model that picks, not one that writes
Before, a text model (llama-3.3-70b-instruct-fp8-fast) got a prompt and had to write back an answer, such as one skill id or NONE. Clef works differently. You send a typed question, either "pick one of these options" or "yes or no", and you get back a probability for every option.
A number per option is what makes the fix work. We can set a floor: if no option is likely enough, nothing fires, and we can log why.
How we measured
- Every quality number below is from an offline test on recorded turns, not live traffic. The samples are small.
- The main set is 141 human turns from the owner's tenants over five days. Replaying the old setup matched production's skill answer on 121 of 121 turns.
- Agreement is not correctness. Where the old and new judges disagreed, a person read the turn and judged it by the shipped rule.
1. The right skill: 4 of 10 became 9 or 10 of 10
On the 10 disputed turns where the rule decides the answer:
| Judge | Right |
|---|---|
| 70B text model (old) | 4 of 10 |
| Clef 27B, same candidates, fire at p ≥ 0.5 | 9 of 10 |
| Clef 27B, whole catalog, fire at p ≥ 0.5 | 10 of 10 |
The bigger change for you is speed at the slow end:
- The slowest 70B call took 42.9 s. In production, a 70B call over 2,000 ms was cut off and logged as "no skill", the same as a real "no skill".
- Now a timeout is logged as a timeout, so a slow call never looks like a real answer.
The floor matters. Clef's top answer fires on one-word test prompts at p 0.3 to 0.4. Without the 0.5 floor, it fired on 21 of 121 turns and was right less often than the old model. With the floor, those fires stop.
Asking over the whole catalog also fixes search misses. The old route first searched for at most 8 candidate skills. On 41 made-up labeled cases, 3 of its 4 misses came from that search finding nothing. The whole-catalog question has no search step and got all 3: 39 of 41 against 37 of 41. That set is small, so it does not prove the two apart on its own.
What shipped: Clef 27B over the whole catalog, fire at p ≥ 0.5, with a hard stop at 2,000 ms.
Limit: these numbers use an 11-skill test catalog, because the owner's tenants have no skills.
2. Memory: fewer facts, and more of them useful
At the start of a conversation, Context Mode sends a short block of facts it remembers about you. Before, it sent about six rows every time, mostly your newest facts, whatever the turn was about.
Now clef-flash reads a pool of 20 rows and asks, for each one, "is this relevant to this turn?" Rows at p ≥ 0.25 ship, best first, at most six.
| On 52 conversation openings | Before | clef-flash filter |
|---|---|---|
| Rows sent that were relevant (precision) | 4.2% | 45.3% |
| Relevant rows that were found (recall) | 27.7% | 61.7% |
| Rows sent | 5.96 | 1.23 |
| Memory tokens added | 262 | 41 |
A 70B judge labeled relevance, twice, in two row orders. On the stricter labels (rows both passes agree on), precision went from 1.6% to 14.1% and recall from 27.8% to 50.0%. Same verdict.
This one adds a step on our side: one extra clef-flash call per opening, with p50 626 ms in the Worker and a hard stop at 1,200 ms. On a timeout or error, the old six rows ship. Your side gets about 200 to 220 fewer input tokens on that opening request. Later turns pay nothing.
3. Resume: every "yes" was wrong
When your first message looks like "continue where we left off", Context Mode replays your last session. A similarity score against 10 sample phrases used to decide this.
- On 87 short turns, the score said yes 12 times. All 7 distinct texts behind those yeses were greetings, approvals, a status question or a remark.
- On the 23 first turns that production actually checks, it said yes twice, both times on a greeting.
- clef-flash said no to all 7 texts, with p from 0.01 to 0.17.
- On made-up phrases, clef-flash gave 0.984 to 0.987 to real resume requests and at most 0.072 to the rest. The old score put "run the tests" (0.639) and "merhaba" (0.604) over the line.
What shipped: when the old score says yes, clef-flash gets one yes/no question, with a hard stop at 300 ms. On a timeout, the old answer stands. The extra call ran on 2 of 141 human turns.
Limit: the test set held no real resume request, so we have not measured how many real ones it catches. We tested Turkish and English only.
4. The one we did not ship
After each turn, a 70B model reads it and pulls out facts to remember. We tried a Clef yes/no question in front of it: "does this turn state a lasting fact?" The best version sent 68% of typed turns to the 70B and lost 1 fact turn of 69. It saved very little.
We left it off. The working threshold sat just above the lowest score Clef gives on this question, so a small model update could change what it catches. And to keep every fact turn, the gate had to send 90% to 99% of turns, so it saved almost nothing. It stays off by default.
5. The biggest saving came from a plain rule
A turn with no words you typed, such as a tool result or a subagent's last turn, can never produce a stored fact. We now skip the 70B call on those turns, with no model involved, and a test runs the real filter to prove nothing is lost.
- On the test set, 51.5% of turns had no typed words, so about half the extraction calls went away.
- On a newer set, 85.6% had none.
- We have not measured the share of such turns across all production traffic, so the real saving sits in that range.
Correctness and speed decided these moves, not price.
Where a decision model helps, and where it does not
- Good at a choice: pick a skill, keep or drop a fact, yes or no on resume.
- Bad at writing: the focus check must return a one-line task summary, so it stays on the 70B, even though Clef spotted task switches more often (16 of 21 against 11 of 21).
- No better when the input is thin: on topic drift, Clef and the old 8B model both scored 10 of 17. Drift stays on the 8B, now with a time limit.
We chose the 0.5 skill floor and the 0.25 memory cut after looking at the data, so each change has its own off switch.
Not measured yet
- Quality on live production traffic. Every quality number above is offline.
- The time cost of the memory filter on a real turn, where it runs beside other reads.
- The skill choice on a real customer catalog.
- Resume detection in languages other than Turkish and English.
How to try it
Run one command:
npx @context-mode/cli
It sets up Claude Code or Codex to send requests through Context Mode. Start free with 1,000 requests, every feature, no card. Your own Anthropic or OpenAI account pays for the model.
- Quick start
- Skills: one catalog, the matching skill loads in any agent.
- Memory: facts shared across your agents and machines.
- Session resume: pick up a past session on any machine.
- Plans and limits
More on the request path these choices run in: Slipstream: a shorter path from your agent to the model.