Slipstream: a shorter path from your agent to the model
Slipstream is the new request path inside Context Mode, live since October 8, 2026 at 02:45 UTC. The memory stores now run in their own small Worker, a request makes at most two store calls before it reaches the model, and memory and skills are picked by newer Cloudflare models. Here is what changed, what we measured, and what went wrong during the rollout.
Where each number comes from
- prod: production traffic.
- A/B: a 2-minute A/B on our proving ground, 203 tenants, the same request sent to the old path and to Slipstream.
- eval: an offline evaluation on recorded turns.
Numbers marked A/B or eval are not yet confirmed on production traffic.
The short version
- An idle memory store answers its first call in 113 ms, down from 394 ms (prod).
- Before the model call, a request makes at most 2 store calls, started together. It used to make 14 to 22, in up to 10 rounds.
- Memory adds 41 tokens per turn instead of 262, and 45% of the facts it adds are relevant, up from 4% (eval).
- The right skill is picked 9 times in 10, up from 4 (eval, small sample).
A shorter path to the model
Before Slipstream, the memory stores lived inside the 11 MB gateway script, and a request read them one step at a time: 14 to 22 calls in up to 10 rounds, each round waiting for the one before it. Slipstream moves the stores into their own 1.4 MB Worker, cm-gateway-state. The data and where it is stored did not change.
Before up to 22 calls, 10 rounds
Slipstream at most 2 calls, 1 round
For 21 of 23 turn types, the gateway now needs at most two store calls before the model call, and it starts both at once. Two more changes take work off the request path:
- Long sessions: only the new messages of a turn are processed. The results for earlier messages are reused instead of computed again on every turn.
- Sign-in: the account lookup moved from a central object on the request path to Workers KV, with an in-memory cache.
What changed
| Area | Before | Now |
|---|---|---|
| Where the memory stores run | Inside the 11 MB gateway script | Their own 1.4 MB Worker; same data, same location |
| Store calls before the model call | 14 to 22, in up to 10 rounds | At most 2, started together (21 of 23 turn types) |
| Long sessions | The whole history processed on every turn | Only the new messages processed |
| Sign-in lookup | A central object on the request path | Workers KV with an in-memory cache |
| Memory recall | bge ranking; all six candidates sent | Cloudflare clef-flash keeps only the relevant ones |
| Skill choice | llama-3.3-70B | Cloudflare Clef 27B |
| Session resume detection | Similarity score only | Similarity score, confirmed by clef-flash |
| Memory extraction | 70B model on every turn | Skipped on turns with no typed user words; nothing is lost |
| Topic drift check | No time limit | A 1.5 s limit |
| Old archives | Kept forever | 7 days on Free, 30 days on Pro |
Speed
| Measure | Before | Now | Source |
|---|---|---|---|
| First call to an idle memory store | 394 ms | 113 ms | prod, Cloudflare measurement, n 36 |
| Gateway time before the model call, p50 | 200 ms | 176 ms | A/B |
| Gateway time before the model call, p90 | 1,138 ms | 905 ms | A/B |
| Cold turn, p50 | 929 ms | 602 ms | A/B |
| Time to the first response header, p90 | 1,876 ms | 1,493 ms | A/B |
Memory and skills
Recall used to rank stored facts with bge and send all six top candidates to the model. Slipstream asks Cloudflare's clef-flash model which of them matter for this turn and sends only those. Fewer facts reach the model, and more of them are right.
| Measure | Before | Now | Source |
|---|---|---|---|
| Recalled facts that were relevant (precision) | 4% | 45% | eval |
| Relevant facts that were found (recall) | 28% | 62% | eval |
| Memory tokens added per turn | 262 | 41 | eval |
| Correct skill choice | 4 of 10 | 9 of 10 | eval, small sample |
| False session resume detections | 9 of 9 | 0 | eval |
| Requests whose memory missed its time budget | 10–16% | 8% | A/B |
| Same, on cold turns | 30–42% | 22–25% | A/B |
| Errors | 0.06–0.86% | 0% | A/B |
A request whose memory misses its time budget still goes to the model on time, without the memory block. That is why the budget misses matter: each one is a turn where the agent did not get the facts it had.
What went wrong during the rollout
- 02:45 to 03:36 UTC, maintenance mode. Agents kept working, with requests passed straight to the model and Context Mode's features off. The cut-over script stopped early on a build status read that later proved wrong; the build itself had succeeded. The remaining steps ran at 08:50.
- 08:33 to 11:04 UTC, Recall reads failed. Turns that use Recall waited up to 15 s. A wrapper called Durable Object RPC methods through
.apply, and on an RPC stub that is itself a call to a remote method namedapply. The fix isReflect.apply, with a regression test. After the fix, 47 of 47 Recall reads succeeded, p50 9 ms. - From 11:00 UTC, store overload under very heavy use. At about 50 to 90 turns a minute on one tenant, its memory store reported "overloaded" on background writes. User requests still succeeded: 500 of 500. The cause is two scans in the store, the archive writer and the insights screen. A fix is in progress.
- Out-of-memory kills. 16 between 10:15 and 10:31 UTC, while Recall reads were stalling. None since the 11:04 fix.
The Recall bug, simplified:
// Before: on an RPC stub, .apply is looked up as a remote method named "apply"
return method.apply(stub, args);
// Now: call the method itself
return Reflect.apply(method, stub, args);
Cleanup
- Removed 20 Workers we no longer need (the proving ground, staging and old experiments) and 13 KV namespaces.
- The proving ground's configs are now generated on deploy, not committed.
Next
A fix for the two store scans behind the overload, and production numbers for the rows marked A/B and eval. Live uptime for every component is on the status page.
What the rebuild taught us about Workers and Durable Objects: Durable Objects in production: 14 lessons from Slipstream.