Repository navigation
Container never sleeps after a client abandons a body-bearing response: inflightRequests leaks and sleepAfter renews forever #242
Description
Activity
We hit this in production and measured it, so here is a reproduction, the cause as far as we can see, and a workaround verified on real containers.
Setup:
@cloudflare/containers0.3.7,compatibility_date2026-09-12,sleepAfter: '10m', one container per app behind Cloudflare Access.What we saw: scanners probing public paths (
GET /assets/.env) gave up during the ~4 s cold start (outcome: clientDisconnectedin Workers observability). After that, the container never slept: 9 h on one app and about 33 h on another, with zero requests, until the Durable Object restarted (a deploy, or aninternal error).A/B on a real app: two sleeping containers of the same app. One was woken by a request read to completion: it slept after about 10 min. The other was woken by a request the client aborted after 1 s: it was still running 29 min later with no traffic, and stayed up for more than 16 h.
Cause, reproduced without a container: a Durable Object that waits 5 s and then returns a body piped through an
IdentityTransformStream, ascontainerFetchdoes:abort reaches the DO? pipeTosettles?client reads the body — yes client aborts at 1 s, default flags no ( request.signalnever aborts)never client aborts at 1 s, enable_request_signalyes (~0.9 s) never, unless something reacts same, and the body is cancelled when request.signal.abortedyes yes (rejects, so .finallyruns)So when the client is gone before the
Responseis returned, nothing ever reads or cancelsreadable, the pipe stays pending,decrementInflight()never runs, andisActivityExpired()keeps renewing the activity timeout.Workaround (in our
Containersubclass, plus theenable_request_signalcompatibility flag):async fetch(request) { const res = await super.fetch(request) if (request.signal?.aborted && res.body) await res.body.cancel('client gone').catch(() => {}) return res }
Confirmed on a real container: woken by a request aborted after 1 s during its cold start, it slept about 11 min later, with
sleepAfterat 10 min. Its sibling without the workaround, woken the same way, stayed awake for more than 16 h with no traffic.Suggested fix in the library: in
containerFetch, also release the inflight slot whenrequest.signalaborts (cancelreadable/ settle the pipe), and document that this needsenable_request_signal. #241 looks like the same leak from the other side (the invocation cancelled whiletcpPort.fetchis pending).Adding a potentially related observation on
@cloudflare/containers0.3.7. We saw the same long-lived idle-instance symptom, but have not confirmed the abandoned-response /inflightRequestsmechanism in our case. Posting here rather than opening another issue because this report and #241 already cover the symptom and a plausible mechanism.Setup
- Worker with
compatibility_date: "2026-10-01", using aContainersubclass withdefaultPort = 8790andsleepAfter = "2m". - A fixed pool of three named instances,
standard-2, with one image-processing request admitted per instance at a time. - Requests are proxied through
getContainer(binding, name).fetch(request)andsuper.fetch(request)in the subclass. - The container runs a Node 24 HTTP service using sharp. It returns JSON containing image derivatives. The Worker is reached through service bindings, without a public route or WebSockets.
Original observation
The affected staging instance was first recorded around 2026-10-08 16:17:34 UTC, during a deployment around 16:17–16:18 UTC; the container application was modified at approximately 16:18:48 UTC. On 2026-10-09, Wrangler still reported that instance as
running, with its recorded creation time more than 20 hours earlier, despitesleepAfter = "2m".A 60-second live-tail observation during the investigation showed no requests. This is a limited observation window, not a complete traffic or continuous-runtime trace for those 20+ hours. Other pool members had been observed sleeping normally.
Initially we suspected a deployment / Durable Object replacement interaction. We do not have the original request-abort history, response-consumption trace, inflight counter, or sufficient alarm evidence to distinguish that hypothesis from the leaks described here and in #241. We cannot claim that an alarm was lost or that either counter-leak mechanism has been reproduced in our application.
Mitigation and verification
We added an independent idle self-exit inside the Node service: exit with code 0 after three minutes with no unfinished work, keeping the Container library's two-minute
sleepAfter. Upload receipt, native derivation, and response completion count as work; client disconnection alone does not mark an unfinished derivation idle.On real staging on 2026-10-09:
- We deployed the new image while all three previous instances were running, stopped test traffic, and verified that every instance became
inactive, including the previously stuck instance. - We started the fixed image and redeployed the Worker while one instance was running, without changing the container image or application version. The deploy log confirmed no container-application changes. The last response completed at 15:39:42.001 UTC; all three instances were observed
inactiveat 15:43:09.486 UTC. - A subsequent cold-start derivation returned HTTP 200 with the expected image output, and all instances were again observed
inactiveafter the full idle window.
These checks validate the application-side backstop; they do not establish that the library issue is fixed or identify which shutdown mechanism ended each process. The timestamps above are control-plane observations, not exact process-exit timestamps. We do not yet have a deterministic reproduction of the original failure without the backstop.
Does this look consistent with the accounting leaks in this issue / #241? If there is specific instrumentation that would distinguish an abandoned response from a deployment lifecycle problem, guidance would help us collect a more decisive reproduction. Account and instance identifiers are omitted from this public comment.
Reacted by Filip Messa- Worker with
Version
@cloudflare/containers0.3.7 (latest at time of filing)Summary
When a caller abandons a body-bearing response from
containerFetch()without consuming the body, the inflight-request counter never decrements,isActivityExpired()renews the activity timeout on every alarm pass, and the container (plus its supervisor DO) runs — and bills — indefinitely.sleepAfternever fires. The only thing that ends the instance is a deploy rolling the image.Where
In
containerFetch()(dist/lib/container.js), the decrement for body-bearing responses is deferred until the body has been fully piped to the consumer:If the returned
readableis never read (client disconnected, caller timed out and abandoned the await, isolate evicted),pipeTostalls on backpressure forever once the body exceeds the transform stream's buffer, so the.finally()never runs.Meanwhile the activity check treats a pinned counter as endless activity:
Small bodies that fit the stream buffer complete the
pipeToregardless of a reader, so this only bites once responses exceed the buffer — which is exactly the large-payload case (file transcodes, archives, model output).How we hit it
An ffmpeg transcode container returns multi-megabyte mp4 bodies to a Cloudflare Workflow; each step does
const res = await container.fetch(...); await res.arrayBuffer(). Any step abandoned between those two awaits (step timeout, retry, deploy mid-run) leaks one inflight slot, and that instance never sleeps again despitesleepAfter = '20m'.Over an 11-day window we accumulated ~1,000 idle instance-hours (memory/disk GiB-seconds ÷ instance size) against ~13 hours of actual vCPU work (~98.7% idle). Instances only stopped on deploys. On
standard-3that idle time was ~$79 of an ~$86 invoice.Expected
An abandoned response body should not keep the container awake indefinitely. Some options:
pipeToto a deadline or to the request's abort signal, so an unread body eventually settles the inflight slot.inflightRequests > 0.Workaround
We subclassed
Containerto track last-request time ourselves and usedschedule()(which the alarm loop processes even in the leaked state) to run a reaper that force-stops (stop(), falling back todestroy()) after 30 minutes with no new requests.