# Durable Objects in production: 14 lessons from Slipstream

October 8, 2026 Engineering 7 min read

Context Mode sits between coding agents and the model provider. On every turn it reads per-tenant state from SQLite-backed Cloudflare Durable Objects (memory, skills, Protect rules and a quota), then forwards the request. In October 2026 we rebuilt that request path as [Slipstream](https://context-mode.com/blog/slipstream). Here are 14 things the rebuild taught us about Workers and Durable Objects. We measured every number below.

## The short version

- Measure an echo Worker before you set a latency budget.
- Keep Durable Object classes in a small script of their own. The first call to an idle object fell from 394 ms to 113 ms.
- One invocation runs at most six object calls at once. Making fewer calls helps more than running them in parallel.
- Never call `.apply`, `.bind` or `.call` on an RPC stub method.
- SQLite skips a partial index unless the query repeats the index's condition.

## 1. Measure the floor before you set a target

Our first plan set latency budgets ("add at most 150 ms") before we knew how much an empty proxy adds. Later we deployed a Worker that only forwards requests, at four placements, and timed it with no model in the loop (at least 60 requests per cell):

| Client in | Edge (no placement) | eu-central-1 | us-east-1 | us-west-1 |
| --- | --- | --- | --- | --- |
| Istanbul | +14 ms | +29 ms | +51 ms | +115 ms |
| US East | +9 ms | +183 ms | +14 ms | +111 ms |
| EU | +7 ms | +8 ms | +48 ms | +120 ms |

**Lesson:** a budget without its floor is a guess. Measure an echo Worker first.

## 2. Put the Worker where your users are, not only where the upstream is

We once moved the gateway to US West, because Smart Placement had picked San Jose and the model API is in the US. Then we looked at who sends the requests: most came from Turkey. The model provider is also served through Cloudflare, so a client reaches it at the nearest edge anyway. US West added about 380 ms for most of our users.

**Lesson:** read your request countries (`cf.country` in Workers Observability) before you choose `placement.region`.

## 3. A Durable Object's location is permanent, so choose the hint deliberately

A `locationHint` counts only on the first `get()` of an object, and an object never moves after that. Warm round trips we measured from a Worker in Frankfurt (FRA): 5 ms to an object created with no hint (it landed in FRA or AMS), 13 ms with `weur`, 19 ms with `eeur`, and 93 to 134 ms to US regions.

**Lesson:** a wrong hint is a permanent tax on every call to that object.

## 4. The slow tail was waking objects up, and our script size set the wake time

An object hibernates after **10 s** with nothing pending. Reading an answer and typing the next prompt takes longer than that, so most turns typed by a person found their objects asleep. In production, the gateway's time before the forward was 87 ms at p50 when the thread had been idle under 5 s, and 551 to 993 ms after 30 s or more.

The objects' own wake code was cheap: constructors took 0.06 to 1.3 ms. The time went somewhere else. Our Durable Object classes lived inside the 11 MB gateway bundle, and waking into a new isolate starts the whole script. We measured it on Cloudflare, with 36 cold calls per arm after 130 to 300 s idle:

**Classes in the 11 MB gateway script** *394 ms* **Classes in their own 1.4 MB script** *113 ms*

First call to an idle object, p50. The difference is 280 ms (95% CI 254 to 308 ms).

**Lesson:** keep Durable Object classes in a small script of their own, and bind them with `script_name`.

## 5. `transferred_classes` moves objects with their data, but the switch takes minutes

We moved six classes from the gateway script to a new state script with a `transferred_classes` migration. In a full-size rehearsal (343,332 rows), we lost 0 rows, no object changed location, and none of the 4,816 calls made during the switch failed.

One surprise: after the transfer deploy, the objects kept running the **old** script's class code for about 300 s, then all switched within 2 s.

**Lesson:** wait at least 5 minutes before you deploy the Worker that depends on the new code.

## 6. One invocation runs at most six object calls at once

The seventh call waits a full round trip, so starting every read in parallel stops helping after six. Our old request path made 14 to 22 calls in up to 10 rounds. Slipstream makes at most 2, started together, on 21 of 23 turn types.

Seven calls two rounds

Forward

Two calls one round

Forward

Each column is one round of calls to Durable Objects, and each dot is one call. The seventh call starts only when one of the first six ends.

**Lesson:** making fewer calls helps more than running them in parallel.

## 7. Never call `.apply`, `.bind` or `.call` on an RPC stub method

This one gave us 2.5 hours of degraded service in production. A Durable Object RPC stub answers *any* property access with another stub, so `stub.get.apply(stub, args)` is a remote call to a method named `apply`:

```
The RPC receiver does not implement the method "apply".
```

We had wrapped stubs in a `Proxy` to count rows per request, and the wrapper called `v.apply(t, a)`. Every Recall read failed, and each affected turn waited out its 15 s budget. Our unit tests passed because they used plain functions. The fix:

```
// in a Proxy get trap over a Durable Object stub
return (...args) => Reflect.apply(v, t, args);   // never v.apply(t, args) or v.bind(t)
```

We found the cause in two deploys of a throwaway Worker bound to the live class.

**Lesson:** reproduce RPC problems against the real class, not a mock.

## 8. SQLite skips a partial index unless the query repeats its condition

Under heavy load (50 to 90 turns a minute on one tenant), a single memory object threw "Durable Object is overloaded. Requests queued for too long." 150 times in 50 minutes. One write path read up to 479,172 rows per call. The table had partial indexes, but the query did not repeat their condition, so SQLite could not prove the index applied and scanned the table:

```
CREATE INDEX … ON arch_scope(nh) WHERE nh <> '';

-- table scan: SQLite cannot prove that nh = ? implies nh <> ''
SELECT … FROM arch_scope WHERE nh = ?;

-- uses the index
SELECT … FROM arch_scope WHERE nh = ? AND nh <> '';
```

Adding `AND nh <> ''` cut that call from 252,121 rows read to 53, and overload errors fell to 0.

**Lesson:** check `EXPLAIN QUERY PLAN` for every hot statement, and watch `cursor.rowsRead`. It tells you both the cost and the latency.

## 9. In a single-threaded object, one slow query slows everyone

An object runs one request at a time. A 4.8 s scan during an archive write made every other route on that object wait for seconds, including reads on the request path.

**Lesson:** keep request-path objects small and their statements bounded, and cache heavy dashboard queries. We cache one for 60 s.

## 10. Use KV for global lookups on every request, not a singleton object

Our credential-to-tenant lookup went through one registry object. Cloudflare's guidance is explicit: don't use one Durable Object as a global singleton, and use Workers KV for configuration read on every request. KV reads measured 3 ms at p50. The registry now writes to KV, and the object is read only for a credential it has never seen.

## 11. Rows written are what you pay for, and deletes and alarms count too

On SQLite-backed Durable Objects, more counts as a written row than you might expect: every index entry on the row, every `DELETE` and every `setAlarm()`. Dropping unused indexes, computing sums on read instead of keeping them up to date with triggers, and writing only when a value changes cut our rows written per request roughly in half.

## 12. Cheaper models fail at extraction, but decision models pass at choosing

We tried five cheaper models to replace llama-3.3-70b for memory extraction. None reached our 90% agreement bar; the best reached 83%. Cloudflare's **Clef** decision models did well where the task is a *choice*. They take typed questions, return a probability for each option, and are billed on input tokens only.

- **Recall filtering with clef-flash:** the share of relevant facts sent went from 4% to 45%, and memory tokens per turn from 262 to 41 (offline eval).
- **Skill choice with Clef 27B:** 9 of 10 correct, against 4 of 10.

Separately, skipping extraction on turns where the user typed nothing lost nothing.

## 13. Prove it in minutes, not hours

Our first A/B replayed real traffic in real time. Each run took 2 to 3 hours, and two tenants were too few to tell signal from machine placement. What worked:

- 203 synthetic tenants at once, six turns each, with one idle gap of 35 to 45 s so the objects sleep;
- the same request sent to both arms at the same moment, and a bootstrap over threads;
- row counters on each request, instead of analytics that arrive 5 to 8 minutes late.

One A/B now takes about 2 minutes, report included.

## 14. Three things that mattered less than we expected

- **Deploy resets:** we deployed 472 times in a week, yet requests near a deploy were no more common among the slowest requests than elsewhere.
- **Parallel agents of one user:** latency did not rise with the number of requests in flight.
- **Object placement within Europe:** a few milliseconds at most, once the Worker is near the user.

## Results

| Measure | Before | After | Source |
| --- | --- | --- | --- |
| First call to an idle object, p50 | 394 ms | 113 ms | prod, measured on Cloudflare, 36 cold calls |
| Gateway time before the forward, p50 | 200 ms | 176 ms | A/B |
| Gateway time before the forward, p90 | 1,138 ms | 905 ms | A/B |
| Cold turn, p50 | 929 ms | 602 ms | A/B |
| Object calls before the forward | 14–22 | ≤ 2 | by design, 21 of 23 turn types |

What changed for users, and what went wrong during the rollout, is in [Slipstream: a shorter path from your agent to the model](https://context-mode.com/blog/slipstream). Live uptime for every component is on the [status page](https://status.context-mode.com).

[Start free](https://context-mode.com/docs/quick-start)[All posts](https://context-mode.com/blog)
