# One night with Cloudflare's on-demand Workers profiling

October 10, 2026 Engineering 7 min read

Every request your coding agent sends through Context Mode passes a Cloudflare Worker and a few Durable Objects. For weeks, our busiest tenant store sometimes answered "Durable Object is overloaded", and requests waited in its queue for seconds. Logs could not tell us why. Then Cloudflare shipped [on-demand profiling for deployed Workers](https://blog.cloudflare.com/workers-on-demand-profiling/). We profiled the live gateway and single Durable Objects, and in one night we fixed what the profiles showed: overload errors went from 29 in 30 minutes to 0.

## The short version

- Profile the one Durable Object that queues, not the whole Worker. That is where we found our causes.
- Turn on source maps before you need them. Minified frames tell you nothing.
- Look for code that is correct but runs too often. Our tests were green the whole time.
- On the busiest store, queue p90 fell from 975 to 2,235 ms to 261 to 908 ms.
- We made mistakes too, including two load tests against production. They are below.

## What the feature does

You ask Cloudflare to profile one deployed Worker version for a few seconds, and you get a standard pprof file back. You can profile the whole Worker, or one Durable Object instance by its namespace id and actor id:

```
# CPU or heap profile of the whole Worker, 10 to 30 s.
cf workers versions profile <version-id> --worker-id <worker> --duration-ms 25000 --profile-type cpu > gw.pb

# One Durable Object instance: namespace id plus the actor id (hex of idFromName).
cf workers versions profile <version-id> --worker-id <worker> \
  --namespace-id <ns-id> --actor-id <actor-id> --duration-ms 25000 --profile-type cpu > store.pb
```

Two things made it useful for us:

- **Source maps.** Our bundle is minified. With `upload_source_maps` on, frames come back as real files and lines (`ledger-store.ts:212`), not `index.js:160914`.
- **Per-object profiles.** A Durable Object is a single-threaded actor. When one instance is slow, every request to it waits. Profiling that one instance is what found our causes.

We read the files with a 100-line Python script, with no Go toolchain. It sums samples by function and by "stack contains", and caps each sample (see the pitfalls).

## The order that worked

1. **Count the errors first.** Group error events by message, Worker and minute in Workers Observability. This tells you which Worker and which minutes to profile.
2. **Find the slow object.** Our admin endpoint reports queue time, rows read and hold time per route for each Durable Object. The busiest tenant store stood out: p90 queue 975 to 2,235 ms, max 5,233 ms.
3. **Profile that object** several times during real traffic, 20 to 25 s each, about a minute apart.
4. **Read each function's share of busy CPU**, not raw time.
5. **Fix, deploy, profile again**, and compare the same numbers.

## What the profiles showed

### The store: correct code that ran on every request

| Function | Share of busy CPU | What it was doing |
| --- | --- | --- |
| Schema setup | 10 to 25% | About 65 CREATE and ALTER statements on every request to two routes. Each ALTER threw "duplicate column" and was caught. |
| Background sweep | 12 to 20% | Listed a whole index in one call, with rows as large as 256 KB. |
| Garbage collector | 3 to 44% | Mostly right after the sweep and large reads. |
| The real work | 12 to 16% | The part that should be there. |

None of this showed in logs. The schema setup was correct. It ran on every request instead of once.

### The meter: parsing state the route never used

`fetch` was 69% of busy CPU. Its first line read and parsed a per-call ring of as much as 100 KB on every request, even on routes that never use it.

### The console: whole tables per call

Our row counters showed console routes that read whole tables: 78,732 rows for the sessions list, 78,638 for daily ops, 41,197 for one chart.

### The gateway: small strings, lots of garbage

In the stateless Worker, the garbage collector took 66 to 78% of busy CPU. The largest JavaScript items: measuring the byte length of the request JSON (15.3%) and hashing it (4.7%), both by building many small strings.

## What we fixed

| Fix | Effect |
| --- | --- |
| Schema statements run once per object instance, with a retry if they throw | 10 to 25% of store CPU to 0 to 1.5% |
| The sweep reads pages of 32 keys, one full pass per hour | 12 to 20% to 0 to 3.9% |
| The meter reads its ring only on the routes that use it | No ring read on other routes |
| Daily ops reads the per-day counters we already keep | 78,638 rows to one day plus at most 30 counter rows |
| Byte length and hash read the JSON value directly | About 20% of gateway CPU on that path removed; forwarded bytes identical |
| Non-urgent calls skip a store for 30 s after it answers "overloaded" | Less pile-up on a busy store |
| A per-isolate cache limited only by count got a 4 MB byte limit | Removes a path to about 100 MB in a 128 MB isolate |
| Background object calls catch, count and log their errors | An overload no longer shows as an unhandled exception |

Each fix has a test that fails on the old code. We checked that the bytes we forward did not change, and ran a real agent session through the gateway before and after each deploy.

**Queue p90, before** *975 to 2,235 ms* **Queue p90, after** *261 to 908 ms*

Busiest tenant store. Bars show the top of each range.

| Busiest store | Before | After |
| --- | --- | --- |
| Schema setup, share of busy CPU | 10 to 25% | 0 to 1.5% |
| Sweep, share | 12 to 20% | 0 to 3.9% |
| Garbage collector, share | 3 to 44% | 4 to 19% |
| Queue p90 | 975 to 2,235 ms | 261 to 908 ms |
| Queue max | 5,233 ms | 341 to 3,328 ms |
| "Durable Object is overloaded" errors | 29 in 30 min | 0 in the hours after |

## Pitfalls we hit

1. **"latest" is not what is running.** It is the newest upload, often a preview with no traffic, and you get "No recent executions were found". Pass the deployed version id.
2. **Space captures about 60 s apart.** Back-to-back captures got `429 Too Many Requests`.
3. **Cap each CPU sample.** A sample carries the time since the previous one, so the first sample after an I/O wait carries the whole wait. Uncapped, a 13 s idle gap looks like a hot function. We cap at 5 ms and report the removed time apart.
4. **A quiet object has no isolate.** "no loaded isolate" means it was idle. Profile during real traffic.
5. **Heap captures of the stateless Worker were often empty.** Two of about ten tries returned samples. CPU captures were reliable.
6. **A heap profile shows what is alive, not what one request makes.** One capture blamed a CSS helper and the UI router for half the heap. Both are built once per isolate and kept. The profile was right; our reading was wrong. Count constructions before you fix what a heap profile shows.
7. **A profile can hold string data.** Keep raw files out of git and share only function names and lines.
8. **A capture that is 95 to 100% garbage collector** tells you there is heap pressure, not where. Pair it with a heap capture.

## What we did wrong

- **We load-tested production.** To prove a memory fix, we sent 20 MB bodies from 15 to 30 test tenants at the live gateway. It reproduced the crash: 2 "exceeded memory limit" errors in the first run, 76 in one minute in the second. A killed isolate fails every request it holds, real users' requests too. We stopped and moved all load proof to local probes. We should have started there.
- **We shipped a fix that was not enough.** The first memory fix bounded a cache, and the crash came back. We should have counted every whole copy of a large body before the first fix.
- **The profile found what our tests did not.** The schema bug had a passing test, because the test checked correctness, not cost.

## What could be better

Our wish list for Cloudflare, from one night of use:

- Reliable heap profiles for stateless Workers. CPU without heap leaves memory kills hard to explain.
- Request size and route on "exceeded memory limit" events. Today we match them by minute.
- A per-object view in the dashboard: which Durable Object instance queued, and for how long.
- Profiling started by an error, for example the next 10 s after an overload, to catch the moment itself.

## Takeaways

1. Turn on source maps before you need them.
2. Profile the one Durable Object that queues, not the whole Worker.
3. Look for correct code that runs too often: schema setup, full scans, parsing state the route does not use.
4. Garbage collector share is a symptom. Find the allocations.
5. Prove load and memory fixes locally. Production profiling is for finding causes, not for stress.

More from the same request path: [Durable Objects in production: 14 lessons from Slipstream](https://context-mode.com/blog/durable-objects-in-production). Live uptime for every component is on the [status page](https://status.context-mode.com). To put your own agent on this path, start with the [quick start](https://context-mode.com/docs/quick-start).

[Start free](https://context-mode.com/docs/quick-start)[All posts](https://context-mode.com/blog)
