Slipstream: a shorter path from your agent to the model

Engineering5 min read

Slipstream is the new request path inside Context Mode, live since October 8, 2026 at 02:45 UTC. The memory stores now run in their own small Worker, a request makes at most two store calls before it reaches the model, and memory and skills are picked by newer Cloudflare models. Here is what changed, what we measured, and what went wrong during the rollout.

Where each number comes from

Numbers marked A/B or eval are not yet confirmed on production traffic.

The short version

A shorter path to the model

Before Slipstream, the memory stores lived inside the 11 MB gateway script, and a request read them one step at a time: 14 to 22 calls in up to 10 rounds, each round waiting for the one before it. Slipstream moves the stores into their own 1.4 MB Worker, cm-gateway-state. The data and where it is stored did not change.

Before up to 22 calls, 10 rounds

Model

Slipstream at most 2 calls, 1 round

Model
Each column is one round of store calls, and each dot is one call. A round starts when the one before it ends. This shows the most calls a request could make, not measured time; the times are in Speed.

For 21 of 23 turn types, the gateway now needs at most two store calls before the model call, and it starts both at once. Two more changes take work off the request path:

What changed

AreaBeforeNow
Where the memory stores runInside the 11 MB gateway scriptTheir own 1.4 MB Worker; same data, same location
Store calls before the model call14 to 22, in up to 10 roundsAt most 2, started together (21 of 23 turn types)
Long sessionsThe whole history processed on every turnOnly the new messages processed
Sign-in lookupA central object on the request pathWorkers KV with an in-memory cache
Memory recallbge ranking; all six candidates sentCloudflare clef-flash keeps only the relevant ones
Skill choicellama-3.3-70BCloudflare Clef 27B
Session resume detectionSimilarity score onlySimilarity score, confirmed by clef-flash
Memory extraction70B model on every turnSkipped on turns with no typed user words; nothing is lost
Topic drift checkNo time limitA 1.5 s limit
Old archivesKept forever7 days on Free, 30 days on Pro

Speed

MeasureBeforeNowSource
First call to an idle memory store394 ms113 msprod, Cloudflare measurement, n 36
Gateway time before the model call, p50200 ms176 msA/B
Gateway time before the model call, p901,138 ms905 msA/B
Cold turn, p50929 ms602 msA/B
Time to the first response header, p901,876 ms1,493 msA/B

Memory and skills

Recall used to rank stored facts with bge and send all six top candidates to the model. Slipstream asks Cloudflare's clef-flash model which of them matter for this turn and sends only those. Fewer facts reach the model, and more of them are right.

MeasureBeforeNowSource
Recalled facts that were relevant (precision)4%45%eval
Relevant facts that were found (recall)28%62%eval
Memory tokens added per turn26241eval
Correct skill choice4 of 109 of 10eval, small sample
False session resume detections9 of 90eval
Requests whose memory missed its time budget10–16%8%A/B
Same, on cold turns30–42%22–25%A/B
Errors0.06–0.86%0%A/B

A request whose memory misses its time budget still goes to the model on time, without the memory block. That is why the budget misses matter: each one is a turn where the agent did not get the facts it had.

What went wrong during the rollout

  1. 02:45 to 03:36 UTC, maintenance mode. Agents kept working, with requests passed straight to the model and Context Mode's features off. The cut-over script stopped early on a build status read that later proved wrong; the build itself had succeeded. The remaining steps ran at 08:50.
  2. 08:33 to 11:04 UTC, Recall reads failed. Turns that use Recall waited up to 15 s. A wrapper called Durable Object RPC methods through .apply, and on an RPC stub that is itself a call to a remote method named apply. The fix is Reflect.apply, with a regression test. After the fix, 47 of 47 Recall reads succeeded, p50 9 ms.
  3. From 11:00 UTC, store overload under very heavy use. At about 50 to 90 turns a minute on one tenant, its memory store reported "overloaded" on background writes. User requests still succeeded: 500 of 500. The cause is two scans in the store, the archive writer and the insights screen. A fix is in progress.
  4. Out-of-memory kills. 16 between 10:15 and 10:31 UTC, while Recall reads were stalling. None since the 11:04 fix.

The Recall bug, simplified:

// Before: on an RPC stub, .apply is looked up as a remote method named "apply"
return method.apply(stub, args);

// Now: call the method itself
return Reflect.apply(method, stub, args);

Cleanup

Next

A fix for the two store scans behind the overload, and production numbers for the rows marked A/B and eval. Live uptime for every component is on the status page.

What the rebuild taught us about Workers and Durable Objects: Durable Objects in production: 14 lessons from Slipstream.