Skip to content

Container never sleeps after a client abandons a body-bearing response: inflightRequests leaks and sleepAfter renews forever #242

Description

@NickBrooks

Version

@cloudflare/containers 0.3.7 (latest at time of filing)

Summary

When a caller abandons a body-bearing response from containerFetch() without consuming the body, the inflight-request counter never decrements, isActivityExpired() renews the activity timeout on every alarm pass, and the container (plus its supervisor DO) runs — and bills — indefinitely. sleepAfter never fires. The only thing that ends the instance is a deploy rolling the image.

Where

In containerFetch() (dist/lib/container.js), the decrement for body-bearing responses is deferred until the body has been fully piped to the consumer:

if (res.body !== null) {
    const { readable, writable } = new IdentityTransformStream();
    res.body?.pipeTo(writable).finally(() => {
        this.decrementInflight();
    });
    return new Response(readable, res);
}

If the returned readable is never read (client disconnected, caller timed out and abandoned the await, isolate evicted), pipeTo stalls on backpressure forever once the body exceeds the transform stream's buffer, so the .finally() never runs.

Meanwhile the activity check treats a pinned counter as endless activity:

isActivityExpired() {
    if (this.inflightRequests > 0) {
        this.renewActivityTimeout();   // renews forever
        return false;
    }
    return this.sleepAfterMs <= Date.now();
}

Small bodies that fit the stream buffer complete the pipeTo regardless of a reader, so this only bites once responses exceed the buffer — which is exactly the large-payload case (file transcodes, archives, model output).

How we hit it

An ffmpeg transcode container returns multi-megabyte mp4 bodies to a Cloudflare Workflow; each step does const res = await container.fetch(...); await res.arrayBuffer(). Any step abandoned between those two awaits (step timeout, retry, deploy mid-run) leaks one inflight slot, and that instance never sleeps again despite sleepAfter = '20m'.

Over an 11-day window we accumulated ~1,000 idle instance-hours (memory/disk GiB-seconds ÷ instance size) against ~13 hours of actual vCPU work (~98.7% idle). Instances only stopped on deploys. On standard-3 that idle time was ~$79 of an ~$86 invoice.

Expected

An abandoned response body should not keep the container awake indefinitely. Some options:

  • Tie the pending pipeTo to a deadline or to the request's abort signal, so an unread body eventually settles the inflight slot.
  • Renew activity only while bytes are actually flowing, rather than unconditionally whenever inflightRequests > 0.
  • At minimum, document that abandoned reads of body-bearing responses pin the container awake, and recommend an application-level watchdog.

Workaround

We subclassed Container to track last-request time ourselves and used schedule() (which the alarm loop processes even in the leaked state) to run a reaper that force-stops (stop(), falling back to destroy()) after 30 minutes with no new requests.

Activity

  1. malheiros commented on Oct 7, 2026

    @malheiros

    We hit this in production and measured it, so here is a reproduction, the cause as far as we can see, and a workaround verified on real containers.

    Setup: @cloudflare/containers 0.3.7, compatibility_date 2026-09-12, sleepAfter: '10m', one container per app behind Cloudflare Access.

    What we saw: scanners probing public paths (GET /assets/.env) gave up during the ~4 s cold start (outcome: clientDisconnected in Workers observability). After that, the container never slept: 9 h on one app and about 33 h on another, with zero requests, until the Durable Object restarted (a deploy, or an internal error).

    A/B on a real app: two sleeping containers of the same app. One was woken by a request read to completion: it slept after about 10 min. The other was woken by a request the client aborted after 1 s: it was still running 29 min later with no traffic, and stayed up for more than 16 h.

    Cause, reproduced without a container: a Durable Object that waits 5 s and then returns a body piped through an IdentityTransformStream, as containerFetch does:

    abort reaches the DO? pipeTo settles?
    client reads the body — yes
    client aborts at 1 s, default flags no (request.signal never aborts) never
    client aborts at 1 s, enable_request_signal yes (~0.9 s) never, unless something reacts
    same, and the body is cancelled when request.signal.aborted yes yes (rejects, so .finally runs)

    So when the client is gone before the Response is returned, nothing ever reads or cancels readable, the pipe stays pending, decrementInflight() never runs, and isActivityExpired() keeps renewing the activity timeout.

    Workaround (in our Container subclass, plus the enable_request_signal compatibility flag):

    async fetch(request) {
      const res = await super.fetch(request)
      if (request.signal?.aborted && res.body) await res.body.cancel('client gone').catch(() => {})
      return res
    }

    Confirmed on a real container: woken by a request aborted after 1 s during its cold start, it slept about 11 min later, with sleepAfter at 10 min. Its sibling without the workaround, woken the same way, stayed awake for more than 16 h with no traffic.

    Suggested fix in the library: in containerFetch, also release the inflight slot when request.signal aborts (cancel readable / settle the pipe), and document that this needs enable_request_signal. #241 looks like the same leak from the other side (the invocation cancelled while tcpPort.fetch is pending).

  2. romansestak commented on Oct 9, 2026

    @romansestak

    Adding a potentially related observation on @cloudflare/containers 0.3.7. We saw the same long-lived idle-instance symptom, but have not confirmed the abandoned-response / inflightRequests mechanism in our case. Posting here rather than opening another issue because this report and #241 already cover the symptom and a plausible mechanism.

    Setup

    • Worker with compatibility_date: "2026-10-01", using a Container subclass with defaultPort = 8790 and sleepAfter = "2m".
    • A fixed pool of three named instances, standard-2, with one image-processing request admitted per instance at a time.
    • Requests are proxied through getContainer(binding, name).fetch(request) and super.fetch(request) in the subclass.
    • The container runs a Node 24 HTTP service using sharp. It returns JSON containing image derivatives. The Worker is reached through service bindings, without a public route or WebSockets.

    Original observation

    The affected staging instance was first recorded around 2026-10-08 16:17:34 UTC, during a deployment around 16:17–16:18 UTC; the container application was modified at approximately 16:18:48 UTC. On 2026-10-09, Wrangler still reported that instance as running, with its recorded creation time more than 20 hours earlier, despite sleepAfter = "2m".

    A 60-second live-tail observation during the investigation showed no requests. This is a limited observation window, not a complete traffic or continuous-runtime trace for those 20+ hours. Other pool members had been observed sleeping normally.

    Initially we suspected a deployment / Durable Object replacement interaction. We do not have the original request-abort history, response-consumption trace, inflight counter, or sufficient alarm evidence to distinguish that hypothesis from the leaks described here and in #241. We cannot claim that an alarm was lost or that either counter-leak mechanism has been reproduced in our application.

    Mitigation and verification

    We added an independent idle self-exit inside the Node service: exit with code 0 after three minutes with no unfinished work, keeping the Container library's two-minute sleepAfter. Upload receipt, native derivation, and response completion count as work; client disconnection alone does not mark an unfinished derivation idle.

    On real staging on 2026-10-09:

    1. We deployed the new image while all three previous instances were running, stopped test traffic, and verified that every instance became inactive, including the previously stuck instance.
    2. We started the fixed image and redeployed the Worker while one instance was running, without changing the container image or application version. The deploy log confirmed no container-application changes. The last response completed at 15:39:42.001 UTC; all three instances were observed inactive at 15:43:09.486 UTC.
    3. A subsequent cold-start derivation returned HTTP 200 with the expected image output, and all instances were again observed inactive after the full idle window.

    These checks validate the application-side backstop; they do not establish that the library issue is fixed or identify which shutdown mechanism ended each process. The timestamps above are control-plane observations, not exact process-exit timestamps. We do not yet have a deterministic reproduction of the original failure without the backstop.

    Does this look consistent with the accounting leaks in this issue / #241? If there is specific instrumentation that would distinguish an abandoned response from a deployment lifecycle problem, guidance would help us collect a more decisive reproduction. Account and instance identifiers are omitted from this public comment.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions