Skip to content

stack: HTTP gateway opens a new upstream connection per request and exhausts ephemeral ports under load #6922

Description

@avallete

Affected area

Local development

Supabase CLI version

develop @ 3704383 (experimental stack runtime)

Operating system

macOS + Docker Desktop

Command

SUPABASE_EXPERIMENTAL_STACK=1 supabase start --runtime docker

Actual output

Under sustained concurrent load, the stack's HTTP gateway (the host-side proxy behind API_URL) returns 502 Bad Gateway in waves. ~/.supabase/stacks/<stack-id>/owner.log shows:

Route rest GET upstream failed before responding, retrying
Route rest request failed { message: "socket hang up", code: "ECONNRESET", _tag: "HttpProxyError", responded: false }

The gateway opens a new upstream TCP connection for every proxied request. Each closed connection sits in TIME_WAIT, so at a few hundred requests per second the ephemeral port range fills up (macOS: 49152–65535). Once it is exhausted, upstream connects are reset and the gateway answers 502 until TIME_WAIT drains, then the cycle repeats.

Expected behavior

The gateway reuses keep-alive connections to each upstream, so sustained load does not exhaust ephemeral ports and requests keep succeeding.

Steps to reproduce

  1. SUPABASE_EXPERIMENTAL_STACK=1 supabase start --runtime docker
  2. Run 24 concurrent client loops for 120 s, each doing GET <API_URL>/rest/v1/ with the service-role key (roughly 480 req/s in total).
  3. Sample netstat -an -p tcp | grep "127.0.0.1.<rest published port> " | grep -c TIME_WAIT every few seconds.

Observed: ~7.3k TIME_WAIT sockets after 20 s and ~15.5k after 40 s, which is the whole ephemeral range. About 64% of requests returned 502, clustered in windows (20–60 s and 100–120 s) that match the TIME_WAIT peaks. Storage routes behave the same way.

Control: hitting the service's published port directly with a keep-alive client at the same load produced zero errors, so the upstream services are fine.

Additional context

Root cause is in packages/stack/src/HttpProxy.ts: forward calls http.request with agent: false, which opens and closes one connection per request.

agent: false was introduced in #6897 because pooled connections could be reset by a just-woken backend (Studio on Docker), and only safe, bodyless requests are retried, so POSTs (e.g. MCP calls) surfaced the reset as a 502. A fix needs to keep that case working:

  • Reuse a pooled keep-alive agent per upstream with a bounded socket count, applied in both native and container runtimes.
  • Handle a reset on a reused socket before any response bytes arrive (the stale keep-alive race) by retrying on a fresh connection, without letting retries multiply new connections.
  • Keep Connection: keep-alive on the wire (Bun ends a streamed upstream response early otherwise).
  • Add an integration test that pushes a few thousand requests through a route and asserts the number of distinct upstream connections accepted by a test server stays bounded.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions