Affected area
Local development
Supabase CLI version
develop @ 3704383 (experimental stack runtime)
Operating system
macOS + Docker Desktop
Command
SUPABASE_EXPERIMENTAL_STACK=1 supabase start --runtime docker
Actual output
Under sustained concurrent load, the stack's HTTP gateway (the host-side proxy behind API_URL) returns 502 Bad Gateway in waves. ~/.supabase/stacks/<stack-id>/owner.log shows:
Route rest GET upstream failed before responding, retrying
Route rest request failed { message: "socket hang up", code: "ECONNRESET", _tag: "HttpProxyError", responded: false }
The gateway opens a new upstream TCP connection for every proxied request. Each closed connection sits in TIME_WAIT, so at a few hundred requests per second the ephemeral port range fills up (macOS: 49152–65535). Once it is exhausted, upstream connects are reset and the gateway answers 502 until TIME_WAIT drains, then the cycle repeats.
Expected behavior
The gateway reuses keep-alive connections to each upstream, so sustained load does not exhaust ephemeral ports and requests keep succeeding.
Steps to reproduce
SUPABASE_EXPERIMENTAL_STACK=1 supabase start --runtime docker
- Run 24 concurrent client loops for 120 s, each doing
GET <API_URL>/rest/v1/ with the service-role key (roughly 480 req/s in total).
- Sample
netstat -an -p tcp | grep "127.0.0.1.<rest published port> " | grep -c TIME_WAIT every few seconds.
Observed: ~7.3k TIME_WAIT sockets after 20 s and ~15.5k after 40 s, which is the whole ephemeral range. About 64% of requests returned 502, clustered in windows (20–60 s and 100–120 s) that match the TIME_WAIT peaks. Storage routes behave the same way.
Control: hitting the service's published port directly with a keep-alive client at the same load produced zero errors, so the upstream services are fine.
Additional context
Root cause is in packages/stack/src/HttpProxy.ts: forward calls http.request with agent: false, which opens and closes one connection per request.
agent: false was introduced in #6897 because pooled connections could be reset by a just-woken backend (Studio on Docker), and only safe, bodyless requests are retried, so POSTs (e.g. MCP calls) surfaced the reset as a 502. A fix needs to keep that case working:
- Reuse a pooled keep-alive agent per upstream with a bounded socket count, applied in both native and container runtimes.
- Handle a reset on a reused socket before any response bytes arrive (the stale keep-alive race) by retrying on a fresh connection, without letting retries multiply new connections.
- Keep
Connection: keep-alive on the wire (Bun ends a streamed upstream response early otherwise).
- Add an integration test that pushes a few thousand requests through a route and asserts the number of distinct upstream connections accepted by a test server stays bounded.
Affected area
Local development
Supabase CLI version
develop@ 3704383 (experimental stack runtime)Operating system
macOS + Docker Desktop
Command
Actual output
Under sustained concurrent load, the stack's HTTP gateway (the host-side proxy behind
API_URL) returns502 Bad Gatewayin waves.~/.supabase/stacks/<stack-id>/owner.logshows:The gateway opens a new upstream TCP connection for every proxied request. Each closed connection sits in
TIME_WAIT, so at a few hundred requests per second the ephemeral port range fills up (macOS: 49152–65535). Once it is exhausted, upstream connects are reset and the gateway answers 502 untilTIME_WAITdrains, then the cycle repeats.Expected behavior
The gateway reuses keep-alive connections to each upstream, so sustained load does not exhaust ephemeral ports and requests keep succeeding.
Steps to reproduce
SUPABASE_EXPERIMENTAL_STACK=1 supabase start --runtime dockerGET <API_URL>/rest/v1/with the service-role key (roughly 480 req/s in total).netstat -an -p tcp | grep "127.0.0.1.<rest published port> " | grep -c TIME_WAITevery few seconds.Observed: ~7.3k
TIME_WAITsockets after 20 s and ~15.5k after 40 s, which is the whole ephemeral range. About 64% of requests returned 502, clustered in windows (20–60 s and 100–120 s) that match theTIME_WAITpeaks. Storage routes behave the same way.Control: hitting the service's published port directly with a keep-alive client at the same load produced zero errors, so the upstream services are fine.
Additional context
Root cause is in
packages/stack/src/HttpProxy.ts:forwardcallshttp.requestwithagent: false, which opens and closes one connection per request.agent: falsewas introduced in #6897 because pooled connections could be reset by a just-woken backend (Studio on Docker), and only safe, bodyless requests are retried, so POSTs (e.g. MCP calls) surfaced the reset as a 502. A fix needs to keep that case working:Connection: keep-aliveon the wire (Bun ends a streamed upstream response early otherwise).