fix(stack): reuse keep-alive upstream connections in the HTTP gateway - #6925
Conversation
The stack gateway opened a new upstream connection for every proxied request. Under sustained load this exhausts ephemeral ports natively (EADDRNOTAVAIL) and makes Docker Desktop's forwarder reset connections, so the gateway answers 502 in waves. Safe requests and requests carrying an Idempotency-Key now go through a keep-alive agent owned by the proxy, with bodies up to 1 MiB buffered so a request that meets a stale pooled connection is replayed once on a fresh one. Other requests keep a fresh connection and are never replayed. Idle pooled sockets close after 4 s, before Node upstreams' keep-alive timeout. Closes #6922 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
Superseded by a newer AI review
🤖 AI Review
Both independent reviews were available. Their findings overlap and merge into one confirmed major retry-safety issue: keyed mutations can execute twice when the upstream does not deduplicate. Verification used the checked-out code, test fixtures, and trusted conventions; tests were not run.
Findings
| Severity | Location | Category | Sources | Claim |
|---|---|---|---|---|
| 🟠 MAJOR | packages/stack/src/HttpProxy.ts:81 |
retry-safety |
claude+codex | Any client-supplied Idempotency-Key enables replay of a buffered mutating request without establishing that the upstream deduplicates it. If a reused connection resets after the mutation executes but before response headers arrive, the proxy sends the mutation again. |
Stats
Claude findings: 1 · Codex findings: 1 · Confirmed: 1 · Refuted: 0 · Uncertain: 0
Models: claude-opus-5-5 + gpt-6.1-sol · Trigger: auto · Workflow run
This review runs once per PR. A maintainer can request another with a /ai-review comment.
Stack upstreams do not deduplicate on Idempotency-Key, so replaying a keyed write after a pooled connection reset could run it twice. Only safe, bodyless requests now share pooled connections and get the single fresh-connection retry; every other request keeps its own connection and is never replayed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
/ai-review |
There was a problem hiding this comment.
🤖 AI Review
Both independent reviews were available. Claude's sole test-clarity nit is refuted: the keyed POST test verifies that an idempotency key does not enable replay. No correctness or security issues were confirmed. Tests were not run.
Findings
No issues found.
Refuted findings (kept for transparency, not posted as review comments)
packages/stack/src/HttpProxy.integration.test.ts:738(test-clarity): The test name and idempotency-key header misleadingly suggest keyed POSTs receive special handling, although pooling and replay depend only on the request method and body.
Refuted: The name accurately describes the exercised request and outcome. HttpProxy.integration.test.ts:762-763 records the complete body before resetting the connection; :774-777 asserts a 502, exactly one recorded body, and only the failure log. These assertions verify that even a keyed POST is not replayed. HttpProxy.ts:330-335 consistently excludes POSTs from pooling and retry. Ignoring the key is the behavior this test protects, not evidence that its name is misleading.
Stats
Claude findings: 1 · Codex findings: 0 · Confirmed: 0 · Refuted: 1 · Uncertain: 0
Models: claude-opus-5-5 + gpt-6.1-sol · Trigger: manual · Workflow run
This review runs once per PR. A maintainer can request another with a /ai-review comment.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
jgoux
left a comment
There was a problem hiding this comment.
Looks good — one test-coverage nit inline.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
TL;DR
The local stack's API gateway no longer opens a new upstream connection for every request. Safe, bodyless requests (GET, HEAD, OPTIONS, TRACE) reuse pooled keep-alive connections, so sustained read traffic stops producing waves of
502 Bad Gateway. Writes keep one connection per request so they are never sent twice.Before
After
flowchart LR C[Client request] --> R{"Safe method<br/>and no body?"} R -- yes --> P["Pooled keep-alive<br/>connection"] R -- no --> F["Fresh connection,<br/>never replayed"] P --> E{"Failed before<br/>a response?"} E -- yes --> O["One retry on a<br/>fresh connection"] E -- no --> D[Response]Why
Since #6897 the gateway forwarded with
agent: false. Under sustained concurrentGET /rest/v1/traffic, closed connections pile up inTIME_WAIT: the native runtime fails withconnect EADDRNOTAVAILonce the ephemeral range is full, and on Docker Desktop the port forwarder starts resetting connections (socket hang up,ECONNRESET). The gateway then answers 502 untilTIME_WAITdrains. Details and measurements are in #6922.#6897 moved away from pooling because Studio's MCP route closes its connection shortly after each POST, so a pooled POST could land on a closing connection and could not be replayed. Those POSTs now stay on fresh connections.
What changed
Closes #6922
🤖 Generated with Claude Code