fix(console,sdk): tear down console/dev children so they cannot spin at 100% CPU - #7297
Conversation
…at 100% CPU Root cause for #6861: simulator sandboxes are forked with detached:true so Ctrl+C does not immediately kill them. Several gaps then left those children orphaned (or the parent hung) while developing the console: 1. Simulator.stop() refused to run while status was "starting", and the console fires simulator.start() without awaiting it — so stopping mid-boot skipped cleanup of already-spawned detached sandboxes. 2. Sandbox.cleanup() only sent SIGTERM and dropped the handle; busy-loop children that ignore SIGTERM survived forever at ~100% CPU. 3. Console HTTP close never terminated the tRPC WebSocketServer, so a promisified server.close() could hang; the app/dev script also had no force-exit timeout and ignored a second Ctrl+C. Fixes: allow stop-during-start with abort + shared stop promise; escalate sandbox cleanup to process-group SIGKILL after a grace period; close WSS clients before the HTTP server; harden scripts/dev.mjs shutdown (timeout + second-signal force exit). Regression tests cover sandbox SIGKILL escalation and stop-while-starting with a long-running Service. Fixes #6861
|
Thanks for opening this pull request! 🎉 |
Adversarial review of #7297 found that the SDK's new "no handle" branch in stopResource never actually reaped the detached sandbox child — the test only checked the SDK's _running state, not whether the child was gone. This commit closes that gap and tightens a few related corners. - simulator: in stopResource, when no handle is recorded in state yet, recover the resource via HandleManager.tryFindHandleByPath (new) and call resource.cleanup() so the sandbox is reaped even if its init() never returns. Deallocate the handle and deregister the policy. - simulator: add tryFindHandleByPath on HandleManager for the inverse handle-by-path lookup. - simulator: in startResource, bail out of writing to state[path].attrs if a stop is in progress between the init() and save() awaits, so a concurrent stopResource doesn't leave a TypeError in its wake. - simulator: reset _stopRequested defensively at the top of update(), so a leaked flag from a previous aborted start cannot silently skip starting resources on hot-reload. - sandbox: in onChildError, use killProcessTree (group kill) for consistency with cleanup(), so grandchildren are reaped too. - dev.mjs: drop optional chaining on forceTimer.unref() — setTimeout always returns a Timeout. - test/simulator/stop-while-starting: handler now writes its PID to a file (path passed via Service.env), and the test asserts process.kill(pid, 0) throws after stop. The previous test would have passed even if the orphan bug persisted. Co-Authored-By: Claude <noreply@anthropic.com>
|
Pushed What
The four low-severity findings (rewording a misleading comment, removing the Pre-existing prettier warnings on |
Summary
Fixes #6861 — Console dev script orphaned node processes at 100% CPU.
Root cause
Detached simulator sandboxes were not reliably cleaned up on stop:
Fix
Test plan