Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 13 additions & 6 deletions src/content/docs/core-concepts.md
Original file line number Diff line number Diff line change
Expand Up @@ -55,7 +55,7 @@ The actual work happens when you call one of the start methods:
| Method | Blocks until | Use when |
| --- | --- | --- |
| `Start()` | `Shutdown` is called or the engine errors | Simplest case; you manage shutdown elsewhere |
| `StartWithContext(ctx)` | `ctx` is cancelled (then graceful shutdown) or the engine errors | You want context-driven lifecycle (signals, parent ctx) |
| `StartWithContext(ctx)` | `ctx` is cancelled and the graceful shutdown it triggers, `OnShutdown` hooks included, has finished; or the engine errors | You want context-driven lifecycle (signals, parent ctx) |
| `StartWithListener(ln)` | as `Start` | Zero-downtime restart via an inherited socket |
| `StartWithListenerAndContext(ctx, ln)` | as `StartWithContext` | Inherited socket + context lifecycle |

Expand Down Expand Up @@ -85,10 +85,17 @@ if err := s.Start(); err != nil {

### Graceful shutdown

`Shutdown(ctx)` stops accepting new connections, drains in-flight requests, then
fires any hooks you registered with `OnShutdown` — in registration order, with
the shutdown context. `StartWithContext` wires this up for you: when the context
is cancelled, the server shuts down using `Config.ShutdownTimeout` (default 30s).
`Shutdown(ctx)` stops the engine, then fires any hooks you registered with
`OnShutdown` — in registration order, with the shutdown context. On `std` and
`adaptive` it waits for in-flight requests before the hooks; on `epoll` and
`io_uring` the engine drains as its listen context is cancelled, and the hooks do
not wait for that (see [Graceful shutdown](/docs/graceful-shutdown#shutdown-sequence)).

`StartWithContext` wires this up for you: when the context is cancelled, the server
shuts down using `Config.ShutdownTimeout` (default 30s), and `StartWithContext`
returns only after that shutdown, hooks included, has finished. A hook must
therefore not wait for `StartWithContext` to return (see
[Graceful shutdown](/docs/graceful-shutdown#drain-hooks-onshutdown)).

```go
ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt)
Expand All @@ -98,7 +105,7 @@ s.OnShutdown(func(ctx context.Context) {
db.Close() // runs during graceful shutdown
})

// Blocks until SIGINT, then drains and runs OnShutdown hooks.
// Blocks until SIGINT; returns after the drain and the OnShutdown hooks.
log.Fatal(s.StartWithContext(ctx))
```

Expand Down
22 changes: 13 additions & 9 deletions src/content/docs/deployment.md
Original file line number Diff line number Diff line change
Expand Up @@ -564,7 +564,7 @@ ready.Store(true) // serving as soon as we're up

s := celeris.New(celeris.Config{Addr: ":8080", ShutdownTimeout: 15 * time.Second})

// Flip readiness to 503 the moment a drain begins, before in-flight requests finish.
// Flip readiness to 503 when Shutdown runs its hooks (see below for when that is).
s.OnShutdown(func(_ context.Context) {
ready.Store(false)
})
Expand All @@ -581,11 +581,14 @@ if err := s.StartWithContext(ctx); err != nil {
}
```

Flipping it in an `OnShutdown` hook (rather than your signal handler) keeps the
readiness change ordered with the rest of the drain. If you prefer, set
`ready.Store(false)` in your own `SIGTERM` handler *before* calling `Shutdown` —
either way the flip is yours to make. (`atomic.Bool` is in the standard library's
`sync/atomic`.)
When that hook runs depends on the engine. On `epoll` and `io_uring` it runs as the
drain begins, while requests are still in flight. On `std` and `adaptive` it runs
only after in-flight requests have finished, which is too late to steer the load
balancer during the drain (see
[Shutdown sequence](/docs/graceful-shutdown#shutdown-sequence)). To flip readiness
before the drain on every engine, set `ready.Store(false)` in your own `SIGTERM`
handler *before* cancelling the context or calling `Shutdown`. Either way the flip is
yours to make. (`atomic.Bool` is in the standard library's `sync/atomic`.)

For true zero-downtime restarts on the same host, inherit the listening socket
across the exec with `InheritListener` + `StartWithListener`
Expand All @@ -607,9 +610,10 @@ drain ordering, and the native engines' `SO_REUSEPORT` rebind — is covered in
[Graceful shutdown and zero-downtime restarts](/docs/graceful-shutdown).

In Kubernetes, the rolling-update pattern is: container receives `SIGTERM` →
your readiness flip fires (the `OnShutdown` hook above) so `/readyz` returns 503 →
LB stops new traffic → in-flight requests drain within `ShutdownTimeout` → process
exits. Set `terminationGracePeriodSeconds` greater than `ShutdownTimeout`.
your readiness flip fires (in your `SIGTERM` handler, or on `epoll` and `io_uring` in
the `OnShutdown` hook above) so `/readyz` returns 503 → LB stops new traffic →
in-flight requests drain within `ShutdownTimeout` → process exits. Set
`terminationGracePeriodSeconds` greater than `ShutdownTimeout`.

## Capacity and timeout tuning

Expand Down
15 changes: 11 additions & 4 deletions src/content/docs/engines.md
Original file line number Diff line number Diff line change
Expand Up @@ -101,8 +101,8 @@ The engines read these at startup. None is needed for normal operation.
| Variable | Engine | Values (default in bold) | Effect |
| -------- | ------ | ------------------------ | ------ |
| `CELERIS_ADAPTIVE_START` | Adaptive | `epoll`, `iouring`, **`auto`** | Chooses the engine Adaptive **starts** on. It does not turn off runtime switching. Unrecognized values mean `auto`. |
| `CELERIS_MAX_IOURING_TIER` | io_uring | `optional`, `high`, `base`, `none` (**unset: detected tier**) | Caps the tier below what the kernel supports; for exercising fallback paths. Any other value, typos included, counts as `none`, and at `none` the io_uring engine reports io_uring as unavailable. |
| `CELERIS_IOURING_SEND_ZC` | io_uring | `on`/`1`/`true`, `off`/`0`/`false`, **`auto`** | Zero-copy send. `auto` enables it where the startup probe finds `SEND_ZC` working; `on` cannot enable it where the probe failed. Unrecognized values mean `auto` and log a warning. |
| `CELERIS_MAX_IOURING_TIER` | io_uring | `optional`, `high`, `base`, `none` (**unset: detected tier**) | Caps the tier below what the kernel supports; for exercising fallback paths. Any other value, typos included, counts as `none`, and at `none` the io_uring engine reports io_uring as unavailable and Adaptive neither starts on io_uring nor switches to it. |
| `CELERIS_IOURING_SEND_ZC` | io_uring | `on`/`1`/`true`, `off`/`0`/`false`, **`auto`** | Zero-copy send. `auto` enables it where the startup probe finds `SEND_ZC` working; `on` cannot enable it where the probe failed. Unrecognized values mean `auto`; one is logged as a warning only where the probe finds `SEND_ZC` working (elsewhere the variable has no effect). |
| `CELERIS_IOURING_MULTISHOT_RECV` | io_uring | `1` (**unset: off**) | Multishot receive into a provided buffer ring (`High` tier). Any value other than `1` leaves it off. |
| `CELERIS_IOURING_PBUF_COUNT` | io_uring | positive integer (**1024**) | Provided-buffer-ring entries per worker; used only with multishot receive. Rounded up to a power of two and clamped to 1024–32768. `0` or an invalid value keeps the default. |
| `CELERIS_IOURING_FIXED_FILES` | io_uring | **do not set** | Development only: fixed-file support is incomplete ([celeris#541](https://github.com/goceleris/celeris/issues/541)), and enabling it makes connections read from unrelated descriptors. |
Expand Down Expand Up @@ -438,7 +438,7 @@ own atomic counters, fetched fresh on each `Metrics()` / `EngineInfo()` call:
| `RequestCount` | `uint64` | Cumulative requests handled by this engine. |
| `ActiveConnections` | `int64` | Currently open connections. |
| `ErrorCount` | `uint64` | Cumulative connection-level or protocol errors. |
| `Throughput` | `float64` | Recent requests-per-second rate. |
| `Throughput` | `float64` | **Always 0**: no engine has ever set it. Deprecated in v1.6.0, removed in v2.0.0 ([celeris#653](https://github.com/goceleris/celeris/issues/653)). Derive a rate from `RequestCount` (example below). |
| `Workers` | `int` | I/O workers (io_uring) or event loops (epoll). Static after `Start`. |
| `AsyncRoutes` | `int` | Count of routes registered `.Async(true)`. Static after `Start`; diagnostics. |
| `AsyncPromotedConns` | `uint64` | Cumulative inline→goroutine promotions via per-handler async. |
Expand All @@ -455,12 +455,19 @@ re-exported on the metrics `Snapshot` as `EngineMetrics`, alongside `RequestsTot
`ErrorsTotal`, `ActiveConns`, `EngineSwitches`, latency buckets, and CPU
utilisation (`celeris/observe/collector.go:40-57`).

A request rate is not one of the counters: take two snapshots and divide the
`RequestCount` difference by the time between them.

```go
const every = 10 * time.Second
prev := s.EngineInfo().Metrics
time.Sleep(every)
m := s.EngineInfo().Metrics
if m.RequestCount > 0 {
rps := float64(m.RequestCount-prev.RequestCount) / every.Seconds()
avgBytes := float64(m.BytesRead+m.BytesWritten) / float64(m.RequestCount)
log.Printf("rps=%.0f conns=%d avg-bytes/req=%.0f promotions=%d",
m.Throughput, m.ActiveConnections, avgBytes, m.AsyncPromotedConns)
rps, m.ActiveConnections, avgBytes, m.AsyncPromotedConns)
}
```

Expand Down
20 changes: 11 additions & 9 deletions src/content/docs/getting-started.md
Original file line number Diff line number Diff line change
Expand Up @@ -188,28 +188,30 @@ func main() {
return c.String(200, "pong")
})

// Blocks until ctx is canceled, then drains in-flight requests
// before returning.
// Blocks until ctx is canceled; returns after in-flight requests
// have drained and the OnShutdown hooks have run.
if err := s.StartWithContext(ctx); err != nil {
log.Fatal(err)
}
}
```

When the context is canceled, Celeris stops accepting new connections and waits
for in-flight requests to finish before returning. The drain window is bounded
by `Config.ShutdownTimeout` (default **30s**). To run cleanup when the server
stops — close a database pool, flush a buffer — register a hook with
`s.OnShutdown`:
When the context is canceled, Celeris stops accepting new connections, and
`StartWithContext` returns only after in-flight requests have finished and your
shutdown hooks have run. The drain window is bounded by `Config.ShutdownTimeout`
(default **30s**). To run cleanup when the server stops — close a database pool,
flush a buffer — register a hook with `s.OnShutdown`:

```go
s.OnShutdown(func(ctx context.Context) {
pool.Close()
})
```

Shutdown hooks fire in registration order with the shutdown context, after the
engine has drained.
Shutdown hooks fire in registration order with the shutdown context. On `std` and
`adaptive` they run after in-flight requests finish; on `epoll` and `io_uring` they
can run while requests are still draining. Either way `StartWithContext` returns only
after they have run, so a hook must not wait for it to return.

> **Tip:** `Config.ShutdownTimeout` only applies to `StartWithContext`. If you
> need a custom drain deadline, set it on the `Config` you pass to
Expand Down
Loading
Loading