Skip to content

std: a cancel of StartWithContext's context ignores Config.ShutdownTimeout; the drain, the OnShutdown hooks and Start wait for every handler #753

Description

@FumingPower3925

Summary

On the std engine, a cancel of StartWithContext's (or StartWithListenerAndContext's) context does not keep to Config.ShutdownTimeout. A handler still running when the budget expires holds the drain, the OnShutdown hooks and the Start call until it returns, however long that is, and the shutdown still reports no error. A direct Server.Shutdown(ctx) on std keeps to its ctx.

Found while measuring, for #738's docs, what a shutdown deadline does to a request still in its handler (lane LIFECYCLE, with the #703 fix, PR #746).

Measured

ShutdownTimeout (or the direct Shutdown's ctx) 500 ms; one request held 2 s in its handler, which either waits for c.Context() or ignores it; linux/arm64 Docker.

Each cell is one run; times are from the moment the shutdown began (738/probe/table.py over the logs):

tree engine shutdown started by handler hooks ran at call that shut down returned client got
main dccb839 std cancel (ShutdownTimeout 500 ms) waits on c.Context() 2123 ms StartWithContext 2123 ms 200 "done" at 2001 ms
main dccb839 std cancel ignores its ctx 2140 ms 2141 ms 200 "done" at 2001 ms
#746 f7685c3, unconstrained memlock std cancel waits on c.Context() / ignores it 2158 / 2161 ms 2158 / 2161 ms 200 "done" at 2000 / 2001 ms
#746 f7685c3, CI shape std cancel waits / ignores 2173 / 2140 ms 2173 / 2140 ms 200 "done" at 2001 / 2000 ms
main and #746 (6 runs) std direct Shutdown(500 ms ctx) either 500-501 ms Shutdown 500-501 ms, context.DeadlineExceeded 200 "done" at 2001-2003 ms
#746 f7685c3 (4 runs) epoll cancel either 500-501 ms StartWithContext at 2000-2001 ms 200 "done" at 2000-2001 ms

In every std cancel run the hook's ctx was already done when it ran, and StartWithContext returned nil. c.Context() was not cancelled in any of the 72 runs (all engines, both modes).

Why

  • The watcher runs Server.shutdown(shutCtx), whose first engine step is std.Engine.Shutdown(shutCtx), a sync.Once around http.Server.Shutdown.
  • But the context handed to Listen is derived from the caller's context, so the cancel reaches std's Listen first: its case <-ctx.Done(): return e.Shutdown(context.Background()) takes the Once with no deadline. The watcher's call then waits in once.Do for that unbounded drain. std's Shutdown comment names this race ("callers race for the once") and relies on the escalation it arms on the caller's ctx: context.AfterFunc(ctx, e.baseCancel) cancels every r.Context() at the deadline (Detached-stream lifecycle follow-ups after #496: H2/h2c SSE disconnect, epoll EPOLLRDHUP skip, std Shutdown, io_uring undrained-send reap #498).
  • The escalation does not reach a celeris handler: Context.Context() returns the stream's context, and on an H1 stream that is context.Background() (stream.(*Stream).Context, h1Mode). Only the SSE/WebSocket detach-close hook is wired to r.Context() (engine/std/bridge.go). So nothing tells the handler, the drain waits for it, and the Once body's error never reaches the watcher, whose Engine.Shutdown call did not run the body and returns nil.
  • A direct Shutdown does not hit this: Listen's context is cancelled only after Engine.Shutdown (the Server.Start / StartWithListener never return after Shutdown on io_uring and epoll: Listen blocks on a Background context and Engine.Shutdown is a no-op #595 ordering), so the caller's ctx wins the Once.

A related doc gap: the engine.Engine interface says of Shutdown "if it expires, remaining connections are closed". No engine does that: in every case above the handler's connection stayed open and the client got the whole response at 2 s.

Fix direction (not tried)

Either keep std's Listen from starting an unbounded drain of its own when a Server.Shutdown with a budget is coming (the watcher always runs one on a cancel since #692), or hand the budget to whichever call wins the Once (for example, Listen waits for the first Shutdown caller's context). Giving handlers r.Context() on std would make the escalation reach a handler that watches c.Context(), but not one that ignores it. The fix belongs to std; epoll, io_uring and adaptive run the hooks at the deadline since #703.

Evidence: evidence/lanes-20260927/LIFECYCLE/738/probe/ in the maintainer's probatorium evidence root (run-probe.sh, deadline_probe_linux_test.go, the logs).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/engineEngine interface or implementationbugSomething isn't working

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions