From 30d1c2d1aaa1a18254d1da477058b7e321c0ee2b Mon Sep 17 00:00:00 2001 From: Albert Bausili Date: Sat, 26 Sep 2026 14:24:54 +0200 Subject: [PATCH 1/4] docs: EngineMetrics.Throughput always reads 0; the parser is not SIMD (celeris#653, celeris#424) - engines.md, observability.md: the Throughput row said "recent requests-per-second rate". No engine has ever set the field, so it always reads 0. It is deprecated in celeris v1.6.0 (goceleris/celeris#695) and removed in v2.0.0 (#651). - engines.md, performance.md: the examples that printed m.Throughput as "rps" now derive the rate from two RequestCount samples. - index.astro: the "SIMD parser" card advertised SSE2/NEON parsing with a SWAR fallback. The only SIMD routine had no caller since celeris 96581bc and is deleted by goceleris/celeris#697. The card now describes the zero-copy parse and keeps the smuggling / rapid-reset hardening, which exists (protocol/h1 framing checks; protocol/h2/stream CVE-2023-44487 mitigation). --- src/content/docs/engines.md | 11 +++++++++-- src/content/docs/observability.md | 2 +- src/content/docs/performance.md | 12 ++++++++++-- src/pages/index.astro | 2 +- 4 files changed, 21 insertions(+), 6 deletions(-) diff --git a/src/content/docs/engines.md b/src/content/docs/engines.md index 0ce9b1f..6746ebb 100644 --- a/src/content/docs/engines.md +++ b/src/content/docs/engines.md @@ -438,7 +438,7 @@ own atomic counters, fetched fresh on each `Metrics()` / `EngineInfo()` call: | `RequestCount` | `uint64` | Cumulative requests handled by this engine. | | `ActiveConnections` | `int64` | Currently open connections. | | `ErrorCount` | `uint64` | Cumulative connection-level or protocol errors. | -| `Throughput` | `float64` | Recent requests-per-second rate. | +| `Throughput` | `float64` | **Always 0**: no engine has ever set it. Deprecated in v1.6.0, removed in v2.0.0 ([celeris#653](https://github.com/goceleris/celeris/issues/653)). Derive a rate from `RequestCount` (example below). | | `Workers` | `int` | I/O workers (io_uring) or event loops (epoll). Static after `Start`. | | `AsyncRoutes` | `int` | Count of routes registered `.Async(true)`. Static after `Start`; diagnostics. | | `AsyncPromotedConns` | `uint64` | Cumulative inline→goroutine promotions via per-handler async. | @@ -455,12 +455,19 @@ re-exported on the metrics `Snapshot` as `EngineMetrics`, alongside `RequestsTot `ErrorsTotal`, `ActiveConns`, `EngineSwitches`, latency buckets, and CPU utilisation (`celeris/observe/collector.go:40-57`). +A request rate is not one of the counters: take two snapshots and divide the +`RequestCount` difference by the time between them. + ```go +const every = 10 * time.Second +prev := s.EngineInfo().Metrics +time.Sleep(every) m := s.EngineInfo().Metrics if m.RequestCount > 0 { + rps := float64(m.RequestCount-prev.RequestCount) / every.Seconds() avgBytes := float64(m.BytesRead+m.BytesWritten) / float64(m.RequestCount) log.Printf("rps=%.0f conns=%d avg-bytes/req=%.0f promotions=%d", - m.Throughput, m.ActiveConnections, avgBytes, m.AsyncPromotedConns) + rps, m.ActiveConnections, avgBytes, m.AsyncPromotedConns) } ``` diff --git a/src/content/docs/observability.md b/src/content/docs/observability.md index 9e821ae..79c4d14 100644 --- a/src/content/docs/observability.md +++ b/src/content/docs/observability.md @@ -130,7 +130,7 @@ counters the adaptive controller reads to pick an engine. | `RequestCount` | `uint64` | Cumulative requests handled by the engine. | | `ActiveConnections` | `int64` | Currently open connections. | | `ErrorCount` | `uint64` | Cumulative connection/protocol-level errors. | -| `Throughput` | `float64` | Recent requests-per-second rate. | +| `Throughput` | `float64` | **Always 0**: no engine has ever set it. Deprecated in v1.6.0, removed in v2.0.0 ([celeris#653](https://github.com/goceleris/celeris/issues/653)). Derive a rate from two `RequestCount` samples. | | `AsyncRoutes` | `int` | Routes registered with `.Async(true)`. | | `AsyncPromotedConns` | `uint64` | Connections promoted to the per-conn dispatch goroutine. | | `Workers` | `int` | Number of I/O workers / event loops. | diff --git a/src/content/docs/performance.md b/src/content/docs/performance.md index d780613..14690ca 100644 --- a/src/content/docs/performance.md +++ b/src/content/docs/performance.md @@ -624,17 +624,25 @@ The latency buckets use fixed bounds of 1 ms, 5 ms, 10 ms, 25 ms, 50 ms, 100 ms, of requests in the high buckets to watch your tail without a full histogram backend. `snap.EngineMetrics` carries the engine-level counters that drive tuning decisions -(`celeris/engine/engine.go:85-132`): `Throughput` (recent RPS), `ActiveConnections`, +(`celeris/engine/engine.go:85-132`): `RequestCount` (sample it twice for a rate), `ActiveConnections`, `AcceptCount` / `CloseCount` (a high close-to-accept ratio means short-lived churn connections), `BytesRead` / `BytesWritten` (the bytes-per-request signal), `Workers`, and the `AsyncRoutes` / `AsyncPromotedConns` dispatch counters from earlier. ```go +const every = 10 * time.Second +prev := s.Collector().Snapshot().EngineMetrics +time.Sleep(every) m := s.Collector().Snapshot().EngineMetrics +rps := float64(m.RequestCount-prev.RequestCount) / every.Seconds() log.Printf("rps=%.0f conns=%d accepts=%d closes=%d async_promotions=%d", - m.Throughput, m.ActiveConnections, m.AcceptCount, m.CloseCount, m.AsyncPromotedConns) + rps, m.ActiveConnections, m.AcceptCount, m.CloseCount, m.AsyncPromotedConns) ``` +`EngineMetrics.Throughput` is not a rate: no engine has ever set it, so it always reads 0. +It is deprecated in v1.6.0 and removed in v2.0.0 +([celeris#653](https://github.com/goceleris/celeris/issues/653)). + If you'd rather not poll, the built-in collector stays on by default (`DisableMetrics: false`) and you can pair it with the in-tree `middleware/metrics` (Prometheus) and `middleware/debug` packages — see diff --git a/src/pages/index.astro b/src/pages/index.astro index aea0a7e..603ce88 100644 --- a/src/pages/index.astro +++ b/src/pages/index.astro @@ -72,7 +72,7 @@ const engines = [ const features = [ { t: "Zero-allocation hot path", d: "Pooled contexts, pre-encoded HPACK, inline param & handler buffers. No interface boxing on the critical path.", icon: '' }, { t: "HTTP/1.1 + h2c", d: "RFC-9112 H1 and full HTTP/2 cleartext with stream multiplexing, flow control and a zero-alloc HEADERS fast path.", icon: '' }, - { t: "SIMD parser", d: "SSE2 (amd64) and NEON (arm64) request parsing with a SWAR fallback — and smuggling / rapid-reset hardening built in.", icon: '' }, + { t: "Hardened zero-copy parser", d: "HTTP/1.1 headers and bodies are parsed in place in the read buffer, not copied out of it — with request-smuggling and HTTP/2 rapid-reset hardening built in.", icon: '' }, { t: "Batteries-included middleware", d: "Auth, CORS, CSRF, rate-limit, circuit-breaker, compress, cache, metrics, OTel, WebSocket, SSE — production-ready, in-tree.", icon: '' }, { t: "Familiar API", d: "Route groups, hierarchical middleware, named routes and error-returning handlers. The Gin/Echo model you already know.", icon: '' }, { t: "Graceful everything", d: "Zero-downtime restart via socket inheritance, draining shutdown, load-shedding overload control and connection hooks.", icon: '' }, From 7e0ca447bffdf7a8e359525b45e389fb4dfb80c2 Mon Sep 17 00:00:00 2001 From: Albert Bausili Date: Sat, 26 Sep 2026 18:18:59 +0200 Subject: [PATCH 2/4] docs: the tier cap at none also stops Adaptive; exact SEND_ZC warning and parser claims (celeris#679, celeris#424) - engines.md, CELERIS_MAX_IOURING_TIER: at `none` Adaptive neither starts on io_uring nor switches to it (goceleris/celeris#694, celeris#679), as the celeris README row now says. - engines.md, CELERIS_IOURING_SEND_ZC: an unrecognized value is logged only where the startup probe finds SEND_ZC working; elsewhere the variable has no effect (engine/iouring/engine.go reads it only inside `if profile.SendZC`, and resolveSendZCPolicy ignores it when the functional probe failed). - index.astro, the parser card: claim what the HTTP/1.1 parser does (its slices alias the bytes it parses), not that every request is parsed in place in the read buffer. epoll and io_uring serve a request that spans reads, has a chunked body, or runs on an async handler from a per-connection buffer. --- src/content/docs/engines.md | 4 ++-- src/pages/index.astro | 2 +- 2 files changed, 3 insertions(+), 3 deletions(-) diff --git a/src/content/docs/engines.md b/src/content/docs/engines.md index 6746ebb..e3249ed 100644 --- a/src/content/docs/engines.md +++ b/src/content/docs/engines.md @@ -101,8 +101,8 @@ The engines read these at startup. None is needed for normal operation. | Variable | Engine | Values (default in bold) | Effect | | -------- | ------ | ------------------------ | ------ | | `CELERIS_ADAPTIVE_START` | Adaptive | `epoll`, `iouring`, **`auto`** | Chooses the engine Adaptive **starts** on. It does not turn off runtime switching. Unrecognized values mean `auto`. | -| `CELERIS_MAX_IOURING_TIER` | io_uring | `optional`, `high`, `base`, `none` (**unset: detected tier**) | Caps the tier below what the kernel supports; for exercising fallback paths. Any other value, typos included, counts as `none`, and at `none` the io_uring engine reports io_uring as unavailable. | -| `CELERIS_IOURING_SEND_ZC` | io_uring | `on`/`1`/`true`, `off`/`0`/`false`, **`auto`** | Zero-copy send. `auto` enables it where the startup probe finds `SEND_ZC` working; `on` cannot enable it where the probe failed. Unrecognized values mean `auto` and log a warning. | +| `CELERIS_MAX_IOURING_TIER` | io_uring | `optional`, `high`, `base`, `none` (**unset: detected tier**) | Caps the tier below what the kernel supports; for exercising fallback paths. Any other value, typos included, counts as `none`, and at `none` the io_uring engine reports io_uring as unavailable and Adaptive neither starts on io_uring nor switches to it. | +| `CELERIS_IOURING_SEND_ZC` | io_uring | `on`/`1`/`true`, `off`/`0`/`false`, **`auto`** | Zero-copy send. `auto` enables it where the startup probe finds `SEND_ZC` working; `on` cannot enable it where the probe failed. Unrecognized values mean `auto`; one is logged as a warning only where the probe finds `SEND_ZC` working (elsewhere the variable has no effect). | | `CELERIS_IOURING_MULTISHOT_RECV` | io_uring | `1` (**unset: off**) | Multishot receive into a provided buffer ring (`High` tier). Any value other than `1` leaves it off. | | `CELERIS_IOURING_PBUF_COUNT` | io_uring | positive integer (**1024**) | Provided-buffer-ring entries per worker; used only with multishot receive. Rounded up to a power of two and clamped to 1024–32768. `0` or an invalid value keeps the default. | | `CELERIS_IOURING_FIXED_FILES` | io_uring | **do not set** | Development only: fixed-file support is incomplete ([celeris#541](https://github.com/goceleris/celeris/issues/541)), and enabling it makes connections read from unrelated descriptors. | diff --git a/src/pages/index.astro b/src/pages/index.astro index 603ce88..1aac246 100644 --- a/src/pages/index.astro +++ b/src/pages/index.astro @@ -72,7 +72,7 @@ const engines = [ const features = [ { t: "Zero-allocation hot path", d: "Pooled contexts, pre-encoded HPACK, inline param & handler buffers. No interface boxing on the critical path.", icon: '' }, { t: "HTTP/1.1 + h2c", d: "RFC-9112 H1 and full HTTP/2 cleartext with stream multiplexing, flow control and a zero-alloc HEADERS fast path.", icon: '' }, - { t: "Hardened zero-copy parser", d: "HTTP/1.1 headers and bodies are parsed in place in the read buffer, not copied out of it — with request-smuggling and HTTP/2 rapid-reset hardening built in.", icon: '' }, + { t: "Hardened zero-copy parser", d: "The HTTP/1.1 parser returns header and body slices that alias the bytes it parses instead of copying them — with request-smuggling and HTTP/2 rapid-reset hardening built in.", icon: '' }, { t: "Batteries-included middleware", d: "Auth, CORS, CSRF, rate-limit, circuit-breaker, compress, cache, metrics, OTel, WebSocket, SSE — production-ready, in-tree.", icon: '' }, { t: "Familiar API", d: "Route groups, hierarchical middleware, named routes and error-returning handlers. The Gin/Echo model you already know.", icon: '' }, { t: "Graceful everything", d: "Zero-downtime restart via socket inheritance, draining shutdown, load-shedding overload control and connection hooks.", icon: '' }, From 15a75d28f97809230c6f418f811173bda0c932d7 Mon Sep 17 00:00:00 2001 From: Albert Bausili Date: Sat, 26 Sep 2026 19:14:53 +0200 Subject: [PATCH 3/4] docs: StartWithContext returns after the OnShutdown hooks; when hooks run vs the drain (celeris#673) celeris#692 (celeris#673) makes a cancelled StartWithContext / StartWithListenerAndContext return only after the Shutdown the cancel triggers has finished, OnShutdown hooks included. graceful-shutdown.md said the opposite ("is not awaited by the return"). It now quotes the new StartWithContext and OnShutdown godoc, says a hook must not wait for StartWithContext to return, and updates the entry-point tables, the hook rules and the pitfalls. core-concepts.md and getting-started.md say the same. Every other page that ordered OnShutdown hooks after the drain, or before it, was checked on Linux with a 500 ms request in flight when the shutdown began. On std and adaptive the hook ran after the request finished; on epoll and io_uring it ran at once, and a direct Shutdown returned at once. So: - graceful-shutdown.md: the Shutdown sequence gains the listen-context step and says which engines Shutdown waits for; the "Shutting down programmatically" paragraph, the Shutdown table row, the shared-budget note, the Drain hooks intro, the socket-handoff step 4 and the FAQ follow. - core-concepts.md, getting-started.md, testing.md: no unconditional "drains, then fires hooks". - deployment.md: the readiness flip in an OnShutdown hook happens before the drain only on epoll and io_uring; on std and adaptive it happens after in-flight requests finish, so flip it in the SIGTERM handler to cover every engine. - The Start() note: since v1.6.0 (celeris#595) Shutdown cancels the listen context of every entry point, so it does stop a server started with Start(). --- src/content/docs/core-concepts.md | 19 ++- src/content/docs/deployment.md | 22 ++-- src/content/docs/getting-started.md | 20 +-- src/content/docs/graceful-shutdown.md | 170 +++++++++++++++++--------- src/content/docs/testing.md | 7 +- 5 files changed, 156 insertions(+), 82 deletions(-) diff --git a/src/content/docs/core-concepts.md b/src/content/docs/core-concepts.md index 3d5ac90..ef16f0e 100644 --- a/src/content/docs/core-concepts.md +++ b/src/content/docs/core-concepts.md @@ -55,7 +55,7 @@ The actual work happens when you call one of the start methods: | Method | Blocks until | Use when | | --- | --- | --- | | `Start()` | `Shutdown` is called or the engine errors | Simplest case; you manage shutdown elsewhere | -| `StartWithContext(ctx)` | `ctx` is cancelled (then graceful shutdown) or the engine errors | You want context-driven lifecycle (signals, parent ctx) | +| `StartWithContext(ctx)` | `ctx` is cancelled and the graceful shutdown it triggers, `OnShutdown` hooks included, has finished; or the engine errors | You want context-driven lifecycle (signals, parent ctx) | | `StartWithListener(ln)` | as `Start` | Zero-downtime restart via an inherited socket | | `StartWithListenerAndContext(ctx, ln)` | as `StartWithContext` | Inherited socket + context lifecycle | @@ -85,10 +85,17 @@ if err := s.Start(); err != nil { ### Graceful shutdown -`Shutdown(ctx)` stops accepting new connections, drains in-flight requests, then -fires any hooks you registered with `OnShutdown` — in registration order, with -the shutdown context. `StartWithContext` wires this up for you: when the context -is cancelled, the server shuts down using `Config.ShutdownTimeout` (default 30s). +`Shutdown(ctx)` stops the engine, then fires any hooks you registered with +`OnShutdown` — in registration order, with the shutdown context. On `std` and +`adaptive` it waits for in-flight requests before the hooks; on `epoll` and +`io_uring` the engine drains as its listen context is cancelled, and the hooks do +not wait for that (see [Graceful shutdown](/docs/graceful-shutdown#shutdown-sequence)). + +`StartWithContext` wires this up for you: when the context is cancelled, the server +shuts down using `Config.ShutdownTimeout` (default 30s), and `StartWithContext` +returns only after that shutdown, hooks included, has finished. A hook must +therefore not wait for `StartWithContext` to return (see +[Graceful shutdown](/docs/graceful-shutdown#drain-hooks-onshutdown)). ```go ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt) @@ -98,7 +105,7 @@ s.OnShutdown(func(ctx context.Context) { db.Close() // runs during graceful shutdown }) -// Blocks until SIGINT, then drains and runs OnShutdown hooks. +// Blocks until SIGINT; returns after the drain and the OnShutdown hooks. log.Fatal(s.StartWithContext(ctx)) ``` diff --git a/src/content/docs/deployment.md b/src/content/docs/deployment.md index 1c863a8..035fbcd 100644 --- a/src/content/docs/deployment.md +++ b/src/content/docs/deployment.md @@ -564,7 +564,7 @@ ready.Store(true) // serving as soon as we're up s := celeris.New(celeris.Config{Addr: ":8080", ShutdownTimeout: 15 * time.Second}) -// Flip readiness to 503 the moment a drain begins, before in-flight requests finish. +// Flip readiness to 503 when Shutdown runs its hooks (see below for when that is). s.OnShutdown(func(_ context.Context) { ready.Store(false) }) @@ -581,11 +581,14 @@ if err := s.StartWithContext(ctx); err != nil { } ``` -Flipping it in an `OnShutdown` hook (rather than your signal handler) keeps the -readiness change ordered with the rest of the drain. If you prefer, set -`ready.Store(false)` in your own `SIGTERM` handler *before* calling `Shutdown` — -either way the flip is yours to make. (`atomic.Bool` is in the standard library's -`sync/atomic`.) +When that hook runs depends on the engine. On `epoll` and `io_uring` it runs as the +drain begins, while requests are still in flight. On `std` and `adaptive` it runs +only after in-flight requests have finished, which is too late to steer the load +balancer during the drain (see +[Shutdown sequence](/docs/graceful-shutdown#shutdown-sequence)). To flip readiness +before the drain on every engine, set `ready.Store(false)` in your own `SIGTERM` +handler *before* cancelling the context or calling `Shutdown`. Either way the flip is +yours to make. (`atomic.Bool` is in the standard library's `sync/atomic`.) For true zero-downtime restarts on the same host, inherit the listening socket across the exec with `InheritListener` + `StartWithListener` @@ -607,9 +610,10 @@ drain ordering, and the native engines' `SO_REUSEPORT` rebind — is covered in [Graceful shutdown and zero-downtime restarts](/docs/graceful-shutdown). In Kubernetes, the rolling-update pattern is: container receives `SIGTERM` → -your readiness flip fires (the `OnShutdown` hook above) so `/readyz` returns 503 → -LB stops new traffic → in-flight requests drain within `ShutdownTimeout` → process -exits. Set `terminationGracePeriodSeconds` greater than `ShutdownTimeout`. +your readiness flip fires (in your `SIGTERM` handler, or on `epoll` and `io_uring` in +the `OnShutdown` hook above) so `/readyz` returns 503 → LB stops new traffic → +in-flight requests drain within `ShutdownTimeout` → process exits. Set +`terminationGracePeriodSeconds` greater than `ShutdownTimeout`. ## Capacity and timeout tuning diff --git a/src/content/docs/getting-started.md b/src/content/docs/getting-started.md index 9f0e626..bc44213 100644 --- a/src/content/docs/getting-started.md +++ b/src/content/docs/getting-started.md @@ -188,19 +188,19 @@ func main() { return c.String(200, "pong") }) - // Blocks until ctx is canceled, then drains in-flight requests - // before returning. + // Blocks until ctx is canceled; returns after in-flight requests + // have drained and the OnShutdown hooks have run. if err := s.StartWithContext(ctx); err != nil { log.Fatal(err) } } ``` -When the context is canceled, Celeris stops accepting new connections and waits -for in-flight requests to finish before returning. The drain window is bounded -by `Config.ShutdownTimeout` (default **30s**). To run cleanup when the server -stops — close a database pool, flush a buffer — register a hook with -`s.OnShutdown`: +When the context is canceled, Celeris stops accepting new connections, and +`StartWithContext` returns only after in-flight requests have finished and your +shutdown hooks have run. The drain window is bounded by `Config.ShutdownTimeout` +(default **30s**). To run cleanup when the server stops — close a database pool, +flush a buffer — register a hook with `s.OnShutdown`: ```go s.OnShutdown(func(ctx context.Context) { @@ -208,8 +208,10 @@ s.OnShutdown(func(ctx context.Context) { }) ``` -Shutdown hooks fire in registration order with the shutdown context, after the -engine has drained. +Shutdown hooks fire in registration order with the shutdown context. On `std` and +`adaptive` they run after in-flight requests finish; on `epoll` and `io_uring` they +can run while requests are still draining. Either way `StartWithContext` returns only +after they have run, so a hook must not wait for it to return. > **Tip:** `Config.ShutdownTimeout` only applies to `StartWithContext`. If you > need a custom drain deadline, set it on the `Config` you pass to diff --git a/src/content/docs/graceful-shutdown.md b/src/content/docs/graceful-shutdown.md index 2e4ab4c..09b19ba 100644 --- a/src/content/docs/graceful-shutdown.md +++ b/src/content/docs/graceful-shutdown.md @@ -20,9 +20,10 @@ listening socket across a restart for true zero-downtime deploys. The idiomatic entry point is `StartWithContext`. Wire a context to your termination signals with the standard library's `signal.NotifyContext`; when the signal lands, -the context is canceled and Celeris drains in flight requests, then `StartWithContext` -returns. This is the recommended path: cancelling the context is what every engine -drains on, so it behaves identically across `std` and the native Linux engines. +the context is canceled, Celeris drains in flight requests and runs your +`OnShutdown` hooks, and only then does `StartWithContext` return. This is the +recommended path: cancelling the context is what every engine drains on, so it +behaves identically across `std` and the native Linux engines. ```go package main @@ -58,25 +59,51 @@ func main() { ``` When `ctx` is canceled, `StartWithContext` cancels the engine's listen context (which -stops accepting and drains in flight requests) and, on a goroutine, calls `Shutdown` -with a fresh context bounded by `Config.ShutdownTimeout` (defaulting to **30s** when -unset or non-positive) to run your `OnShutdown` hooks. `StartWithContext` then returns -the engine's exit error. Source: `celeris/server.go:771-794`. - -> `StartWithContext` is blocking. It returns once the engine has stopped accepting and -> drained, or when the engine returns a fatal error. The internal `Shutdown` goroutine -> (which runs your `OnShutdown` hooks) is not awaited by the return, so if you need to -> be certain a hook finished before the process exits, do that hook's blocking work -> synchronously inside the hook and keep the drain budget large enough — see -> [Drain hooks](#drain-hooks-onshutdown). +stops accepting and drains in flight requests) and, on a watcher goroutine, calls +`Shutdown` with a fresh context bounded by `Config.ShutdownTimeout` (defaulting to +**30s** when unset or non-positive), which runs your `OnShutdown` hooks. It then waits +for both: `StartWithContext` returns the engine's exit error only after the engine has +stopped **and** that `Shutdown`, hooks included, has returned. If your code already +called `Shutdown` itself during the run, the cancel does not run it, or your hooks, a +second time. `StartWithListenerAndContext` behaves the same way. Source: +`celeris/server.go` (`StartWithContext`, `StartWithListenerAndContext`, +`listenUntilCancelled`). + +The godoc states the contract (since celeris v1.6.0, +[celeris#673](https://github.com/goceleris/celeris/issues/673)): + +> **`StartWithContext`:** "When the context is canceled, the server shuts down +> gracefully using Config.ShutdownTimeout (default 30s), and StartWithContext returns +> once that shutdown, including the [Server.OnShutdown] hooks, has finished. A hook +> must therefore not wait for StartWithContext to return; see [Server.OnShutdown]." +> +> **`OnShutdown`:** "When cancelling the context of [Server.StartWithContext] or +> [Server.StartWithListenerAndContext] is what shuts the server down, the hooks run +> before that call returns: it waits for the Shutdown the cancel triggers, hooks +> included. A hook must therefore not wait for that Start call to return, directly or +> through anything that happens only after it returns. The two would wait on each +> other: a hook that returns when its ctx is done ends the wait after +> [Config.ShutdownTimeout], and a hook that ignores ctx never does." + +So a `main` that exits as soon as `StartWithContext` returns does not cut a hook short. +The one thing a hook must not do is wait for `StartWithContext` (or +`StartWithListenerAndContext`) to return, for example by blocking on a channel your +`main` closes after the call returns. A hook that selects on its `ctx` is released +after `Config.ShutdownTimeout`; one that ignores `ctx` hangs the shutdown for good. + +> Before v1.6.0, `StartWithContext` did not wait for the hooks: it could return before +> they had run, and a cancel could skip `Shutdown`, and so the hooks, altogether +> ([celeris#673](https://github.com/goceleris/celeris/issues/673)). ### Shutting down programmatically: `Shutdown` `Shutdown(ctx)` is the explicit, programmatic way to stop a running server. It stops -accepting new connections, drains in flight requests bounded by the `ctx` you pass, -closes the internal CPU monitor, and runs your `OnShutdown` hooks. `Config.ShutdownTimeout` -is **not** consulted on this path — *you* own the deadline via the context you pass. -Source: `celeris/server.go:367-382`. +the engine, closes the internal CPU monitor, and runs your `OnShutdown` hooks. On `std` +and `adaptive` it first waits for in flight requests, bounded by the `ctx` you pass; on +`epoll` and `io_uring` it returns without waiting for them (see +[Shutdown sequence](#shutdown-sequence)). `Config.ShutdownTimeout` is **not** consulted +on this path — *you* own the deadline via the context you pass. Source: +`celeris/server.go` (`Shutdown`). ```go // You own the drain deadline here. @@ -87,64 +114,86 @@ if err := s.Shutdown(shutCtx); err != nil { } ``` -> **Prefer `StartWithContext` over a bare `Start()` + `Shutdown()` pair.** `Start()` -> runs the engine on a non-cancelable background context, so on the native Linux -> engines (`epoll`, `io_uring`) — whose drain is driven entirely by listen-context -> cancellation — a later `Shutdown` call cannot unwind the running engine. The -> context-driven entry points (`StartWithContext` / `StartWithListenerAndContext`) -> cancel the listen context for you and work uniformly on every engine. Source: -> `celeris/server.go:354-360` (`Start` uses `context.Background()`), -> `celeris/engine/epoll/engine.go:157-159`, `celeris/engine/iouring/engine.go:305-309` -> (native `Shutdown` is a no-op; drain happens on listen-context cancellation). +> **Prefer `StartWithContext` over a bare `Start()` + `Shutdown()` pair.** On the +> native Linux engines (`epoll`, `io_uring`) the drain is driven entirely by +> listen-context cancellation: their engine-level `Shutdown` is a no-op. Since v1.6.0 +> ([celeris#595](https://github.com/goceleris/celeris/issues/595)), `Server.Shutdown` +> cancels the listen context of every entry point, `Start()` included, so a later +> `Shutdown` does stop a server started with `Start()`; before v1.6.0, `Start()` ran on +> a non-cancelable background context and `Shutdown` could not unwind a native engine. +> The context-driven entry points (`StartWithContext` / `StartWithListenerAndContext`) +> are still the better path: a cancel runs `Shutdown` for you with +> `Config.ShutdownTimeout`, and the call returns only after the engine and the hooks +> have finished, while a direct `Shutdown` on `epoll` or `io_uring` returns without +> waiting for the drain (see [Shutdown sequence](#shutdown-sequence)). Source: +> `celeris/server.go` (`Start`, `listenContext`, `Shutdown`), +> `celeris/engine/epoll/engine.go` and `celeris/engine/iouring/engine.go` (`Shutdown`). ### Entry points at a glance | Method | Blocks until | Drain deadline | Use when | | ------ | ------------ | -------------- | -------- | -| `StartWithContext(ctx)` | engine drained or engine error | `Config.ShutdownTimeout` (default 30s), applied to the hook phase | The common case: signal-driven shutdown. | -| `StartWithListenerAndContext(ctx, ln)` | engine drained or engine error | `Config.ShutdownTimeout` (default 30s), applied to the hook phase | Socket handoff + signal-driven shutdown. | -| `Start()` | engine error or process exit | n/a (drain via `StartWithContext`) | Rare; prefer the context entry points. | -| `StartWithListener(ln)` | engine error or process exit | n/a (drain via `StartWithListenerAndContext`) | Socket handoff with the context entry point below. | -| `Shutdown(ctx)` | drain + hooks complete | the `ctx` you pass | Programmatic shutdown from your own code. | +| `StartWithContext(ctx)` | after a cancel: the engine has stopped and `Shutdown`, hooks included, has finished; otherwise an engine error | `Config.ShutdownTimeout` (default 30s), applied to the hook phase | The common case: signal-driven shutdown. | +| `StartWithListenerAndContext(ctx, ln)` | as `StartWithContext` | `Config.ShutdownTimeout` (default 30s), applied to the hook phase | Socket handoff + signal-driven shutdown. | +| `Start()` | `Shutdown` is called (since v1.6.0) or engine error | n/a (drain via `StartWithContext`) | Rare; prefer the context entry points. | +| `StartWithListener(ln)` | as `Start()` | n/a (drain via `StartWithListenerAndContext`) | Socket handoff with the context entry point below. | +| `Shutdown(ctx)` | hooks complete; on `std` and `adaptive` also the drain, which `epoll` and `io_uring` do not wait for (see [Shutdown sequence](#shutdown-sequence)) | the `ctx` you pass | Programmatic shutdown from your own code. | Source: `celeris/server.go:354`, `367`, `705`, `716`, `771`. ## Shutdown sequence `Shutdown(ctx)` runs a fixed, well-defined sequence. Knowing the order matters when -you register hooks that depend on it. Source: `celeris/server.go:367-382`. +you register hooks that depend on it. Source: `celeris/server.go` (`Shutdown`), and each +engine's `Shutdown`. 1. **Returns immediately if never started.** If the server was never started (no engine installed), `Shutdown` closes the CPU monitor (a no-op if it was never created) and returns `nil`. Calling `Shutdown` on a server you never started is therefore safe and cheap. -2. **Stop accepting, then drain.** The engine stops accepting new connections and - drains in flight requests, bounded by the `ctx` you pass: per the engine contract, - when the deadline expires remaining connections are closed rather than waited on - indefinitely (`celeris/engine/engine.go:16-18`). -3. **Close the CPU monitor.** Celeris releases the internal CPU-utilization monitor +2. **Shut the engine down.** On `std` this is net/http's `Server.Shutdown`: stop + accepting, then wait for in flight requests, bounded by the `ctx` you pass (per the + engine contract, when the deadline expires remaining connections are closed rather + than waited on indefinitely, `celeris/engine/engine.go:16-18`). `adaptive` waits, + bounded by `ctx`, for its engines to unwind. On `epoll` and `io_uring` this step + returns at once: those engines drain as their listen context is cancelled (the next + step), and `Shutdown` does not wait for that. +3. **Cancel the listen context.** This is what stops a running `epoll` or `io_uring` + engine. `Shutdown` does not wait for it to finish draining. +4. **Close the CPU monitor.** Celeris releases the internal CPU-utilization monitor (on Linux this frees the `/proc/stat` file descriptor that powers the adaptive engine and `CPUUtilization` metrics). -4. **Fire `OnShutdown` hooks.** Your registered hooks run **in registration order**, +5. **Fire `OnShutdown` hooks.** Your registered hooks run **in registration order**, each receiving the **same** shutdown context you passed to `Shutdown`. A panic in one hook is recovered and does not abort the others, nor does it crash the process. -`Shutdown` returns the engine's drain error (or `nil`). Note that hook panics are +So when the hooks run relative to the drain depends on the engine. On `std` and +`adaptive` they start after in flight requests have finished. On `epoll` and `io_uring` +they can run, and a direct `Shutdown` can return, while requests are still in flight. +A cancelled `StartWithContext` still returns only after both the engine and the hooks +have finished. (Measured on Linux with a 500 ms request in flight when the shutdown +began: on `epoll` and `io_uring` the hook ran at once, and a direct `Shutdown` returned +at once; on `std` and `adaptive` the hook ran after the request finished.) + +`Shutdown` returns the engine's shutdown error (or `nil`). Note that hook panics are swallowed (recovered) — they do not surface in the return value — so do your own error logging inside the hook. > **The shutdown context is shared across the engine drain *and* every hook.** Within a > single `Shutdown(ctx)` call, the same `ctx` bounds the drain and then flows into each -> hook in turn — so if you pass a 5s context and the drain eats 4.5s, your hooks have -> only ~500ms of budget left. Size `ShutdownTimeout` (or the context you build manually) -> to cover both the request drain *and* the slowest resource you close in a hook. +> hook in turn — so on `std` and `adaptive`, where `Shutdown` waits for the drain, if +> you pass a 5s context and the drain eats 4.5s, your hooks have only ~500ms of budget +> left. Size `ShutdownTimeout` (or the context you build manually) to cover both the +> request drain *and* the slowest resource you close in a hook. ## Drain hooks: `OnShutdown` -`Server.OnShutdown(fn)` registers a function to run during `Shutdown`, after the -request drain completes. This is where you close database pools, flush log buffers, -deregister from service discovery, or persist in-memory state. Source: -`celeris/server.go:224-227`. +`Server.OnShutdown(fn)` registers a function to run at the end of `Shutdown`. On `std` +and `adaptive` that is after in flight requests have finished; on `epoll` and +`io_uring` a hook can run while requests are still in flight (see +[Shutdown sequence](#shutdown-sequence)). This is where you close database pools, +flush log buffers, deregister from service discovery, or persist in-memory state. +Source: `celeris/server.go` (`OnShutdown`). ```go s := celeris.New(celeris.Config{ @@ -183,7 +232,13 @@ Rules to internalize: - **Respect the context deadline.** The same `ctx` flows into every hook. Long-running cleanup should select on `ctx.Done()` and bail out rather than block the process from exiting. Celeris does **not** forcibly interrupt a hook that ignores the - deadline — it will run to completion and delay your process exit. + deadline — it will run to completion, and a cancelled `StartWithContext` does not + return until it has. +- **Never wait for `StartWithContext` inside a hook.** On the `StartWithContext` / + `StartWithListenerAndContext` path the hooks run *before* that call returns, so a + hook that waits for it to return, directly or on something your code does only + after it returns, waits on itself. It is released after `Config.ShutdownTimeout` if + it honours `ctx`, and never if it does not. - **Panics are contained.** A panic in one hook is recovered; remaining hooks still run and the process does not crash. This is a safety net, not a license to skip error handling — log failures yourself. @@ -426,9 +481,10 @@ The handoff dance (which the parent/supervisor performs) is, in outline: through (the child inherits open fds across `exec`). 3. The child calls `InheritListener("CELERIS_LISTENER_FD")` and starts accepting on the same socket. -4. The parent sends itself (or is sent) `SIGTERM`, drains in flight requests via its - `StartWithContext`/`StartWithListenerAndContext` cancellation, runs its `OnShutdown` - hooks, and exits. +4. The parent sends itself (or is sent) `SIGTERM`. Its + `StartWithContext`/`StartWithListenerAndContext` cancellation drains in flight + requests and runs its `OnShutdown` hooks, and the call returns once both are done, + so the parent then exits. During steps 3–4 both processes accept on the port (native engines via `SO_REUSEPORT`, std via the shared inherited fd), so no client connection is refused. @@ -446,9 +502,13 @@ std via the shared inherited fd), so no client connection is refused. (or match the listener exactly). - **Registering `OnShutdown` after `Start`.** Hooks (and all configuration) must be registered before the server starts. -- **Under-sizing the shutdown budget.** `ShutdownTimeout` (or your manual context) - covers the request drain *and* every `OnShutdown` hook, sharing one deadline. If your - hooks do real work (flushing a remote sink, closing pools), budget for it. +- **Waiting for `StartWithContext` inside a hook.** The hooks run before a cancelled + `StartWithContext` returns, so the two wait on each other. See + [Drain hooks](#drain-hooks-onshutdown). +- **Under-sizing the shutdown budget.** On `std` and `adaptive`, `ShutdownTimeout` (or + your manual context) covers the request drain *and* every `OnShutdown` hook, sharing + one deadline. If your hooks do real work (flushing a remote sink, closing pools), + budget for it. - **Expecting hook panics in the return value.** Hook panics are recovered and *not* reflected in `Shutdown`'s return. Log errors inside the hook. - **Using `PauseAccept` on the std engine.** It returns @@ -468,7 +528,7 @@ Yes. It returns `nil` immediately (after a harmless CPU-monitor cleanup). Source **Do `OnShutdown` hooks run if the engine never started?** No. If no engine was installed, `Shutdown` returns before reaching the hook loop. Hooks -fire only after a real drain. +fire only when an engine was started. **What happens to requests still running when the deadline expires?** Per the engine contract, the engine closes remaining connections rather than waiting diff --git a/src/content/docs/testing.md b/src/content/docs/testing.md index 40566d4..2e52400 100644 --- a/src/content/docs/testing.md +++ b/src/content/docs/testing.md @@ -542,9 +542,10 @@ Key APIs in play: (`celeris/server.go:434`). - **`s.Start() error`** — runs the accept loop; it blocks, so call it in a goroutine (`celeris/server.go:354`). -- **`s.Shutdown(ctx) error`** — stops accepting new connections, drains in-flight - requests, fires `OnShutdown` hooks, and returns `nil` if the server was never - started. Always give it a bounded context (`celeris/server.go:367`). +- **`s.Shutdown(ctx) error`** — stops the engine and fires `OnShutdown` hooks, and + returns `nil` if the server was never started. On `std` and `adaptive` it waits for + in-flight requests; on `epoll` and `io_uring` it returns without waiting for them. + Always give it a bounded context (`celeris/server.go:367`). > Routes must be registered **before** `Start` — handler chains are baked at > registration time, and the `*Server` is only safe for concurrent use after From 155de887d232f27192a6432dbd45b7aaf17db899 Mon Sep 17 00:00:00 2001 From: Albert Bausili Date: Sun, 27 Sep 2026 18:29:06 +0200 Subject: [PATCH 4/4] docs: a Shutdown after a cancel has started one waits for it and keeps its ShutdownTimeout (celeris#673, as merged in #692) --- src/content/docs/graceful-shutdown.md | 35 +++++++++++++++++++++------ 1 file changed, 27 insertions(+), 8 deletions(-) diff --git a/src/content/docs/graceful-shutdown.md b/src/content/docs/graceful-shutdown.md index 09b19ba..a9fcdd6 100644 --- a/src/content/docs/graceful-shutdown.md +++ b/src/content/docs/graceful-shutdown.md @@ -65,7 +65,10 @@ stops accepting and drains in flight requests) and, on a watcher goroutine, call for both: `StartWithContext` returns the engine's exit error only after the engine has stopped **and** that `Shutdown`, hooks included, has returned. If your code already called `Shutdown` itself during the run, the cancel does not run it, or your hooks, a -second time. `StartWithListenerAndContext` behaves the same way. Source: +second time; and a `Shutdown` your code calls after the cancel has started one waits +for that one instead of running its own (see +[Shutting down programmatically](#shutting-down-programmatically-shutdown)). +`StartWithListenerAndContext` behaves the same way. Source: `celeris/server.go` (`StartWithContext`, `StartWithListenerAndContext`, `listenUntilCancelled`). @@ -101,9 +104,23 @@ after `Config.ShutdownTimeout`; one that ignores `ctx` hangs the shutdown for go the engine, closes the internal CPU monitor, and runs your `OnShutdown` hooks. On `std` and `adaptive` it first waits for in flight requests, bounded by the `ctx` you pass; on `epoll` and `io_uring` it returns without waiting for them (see -[Shutdown sequence](#shutdown-sequence)). `Config.ShutdownTimeout` is **not** consulted -on this path — *you* own the deadline via the context you pass. Source: -`celeris/server.go` (`Shutdown`). +[Shutdown sequence](#shutdown-sequence)). When your call is what shuts the server down, +`Config.ShutdownTimeout` is **not** consulted — *you* own the deadline via the context +you pass. Source: `celeris/server.go` (`Shutdown`). + +The exception is a call made after cancelling the context of `StartWithContext` or +`StartWithListenerAndContext` has already started a shutdown. Since celeris v1.6.0 +([celeris#673](https://github.com/goceleris/celeris/issues/673)), that call does not run +a second shutdown, which would run your hooks again. The `Shutdown` godoc: + +> "When cancelling the context of [Server.StartWithContext] or +> [Server.StartWithListenerAndContext] has already started a shutdown, Shutdown does +> not run a second one: it waits for that one, hooks included, and returns its result, +> or ctx's error if ctx is done first." + +That shutdown keeps its `Config.ShutdownTimeout` deadline. The `ctx` you pass bounds +only how long your call waits for it; if `ctx` is done first, the shutdown carries on +without your call. ```go // You own the drain deadline here. @@ -137,7 +154,7 @@ if err := s.Shutdown(shutCtx); err != nil { | `StartWithListenerAndContext(ctx, ln)` | as `StartWithContext` | `Config.ShutdownTimeout` (default 30s), applied to the hook phase | Socket handoff + signal-driven shutdown. | | `Start()` | `Shutdown` is called (since v1.6.0) or engine error | n/a (drain via `StartWithContext`) | Rare; prefer the context entry points. | | `StartWithListener(ln)` | as `Start()` | n/a (drain via `StartWithListenerAndContext`) | Socket handoff with the context entry point below. | -| `Shutdown(ctx)` | hooks complete; on `std` and `adaptive` also the drain, which `epoll` and `io_uring` do not wait for (see [Shutdown sequence](#shutdown-sequence)) | the `ctx` you pass | Programmatic shutdown from your own code. | +| `Shutdown(ctx)` | hooks complete; on `std` and `adaptive` also the drain, which `epoll` and `io_uring` do not wait for (see [Shutdown sequence](#shutdown-sequence)) | the `ctx` you pass, unless a cancel has already started the shutdown (see [above](#shutting-down-programmatically-shutdown)) | Programmatic shutdown from your own code. | Source: `celeris/server.go:354`, `367`, `705`, `716`, `771`. @@ -518,9 +535,11 @@ std via the shared inherited fd), so no client connection is refused. **What's the default drain timeout?** 30 seconds — used by `StartWithContext` and `StartWithListenerAndContext` when -`Config.ShutdownTimeout` is zero or negative. When you call `Shutdown(ctx)` yourself -there is no default; you supply the context (`celeris/config.go:109-111`, -`celeris/server.go:777-780`). +`Config.ShutdownTimeout` is zero or negative. When your own `Shutdown(ctx)` call is +what shuts the server down there is no default; you supply the context +(`celeris/config.go:109-111`, `celeris/server.go:777-780`). A call made after a cancel +has started the shutdown waits for that one, which keeps its `Config.ShutdownTimeout` +deadline. **Is calling `Shutdown` on a server I never started safe?** Yes. It returns `nil` immediately (after a harmless CPU-monitor cleanup). Source: