From 014544df3310196b54ea503609c756159560ce99 Mon Sep 17 00:00:00 2001 From: Nimrod Teich Date: Sun, 23 Aug 2026 11:59:28 +0300 Subject: [PATCH 1/2] docs: document the fleet tracker gate behind --shared-state (MAG-2983) smart-router PR #322 (MAG-2981) makes --shared-state share per-endpoint chain-tracker poll observations across router replicas through the cache, so an upstream is polled about once per interval fleet-wide instead of once per replica, and adds rpc_endpoint_tracker_gate_skips_total{source}. - cli.md: --shared-state row now names both things it shares; a short paragraph explains the borrowing rule and its three safety floors; the Polling relief section gains the flag as the fleet-level lever and points at the gate-skips counter. - metrics.md: new counter row, the quick-answer pointer, and a one-line recipe for confirming the gate is working across replicas. - cache.md: "Sharing state across routers" lists both shared things. --- docs/deployment/cache.md | 14 ++++++++++++-- docs/reference/cli.md | 26 ++++++++++++++++++++++---- docs/reference/metrics.md | 9 ++++++++- 3 files changed, 42 insertions(+), 7 deletions(-) diff --git a/docs/deployment/cache.md b/docs/deployment/cache.md index e0f80fd..cbbeb76 100644 --- a/docs/deployment/cache.md +++ b/docs/deployment/cache.md @@ -99,8 +99,18 @@ flags as the router. ## Sharing state across routers When several router instances share one cache, add `--shared-state` to the routers so -they also share consumer-consistency state through it — keeping their "seen block" views -aligned. See the [CLI reference](../reference/cli.md#cache-shared-state). +they also share two things through it: + +- **consumer-consistency state** — their "seen block" views stay aligned, so a client + that just read block N from one replica is not served N-1 by another; +- **chain-tracker poll observations** — each replica publishes the upstream polls it + makes and borrows fresh ones from the others, so every upstream is polled about + once per interval fleet-wide instead of once per replica. This is the lever that + stops the router's own polling load from growing with the replica count. + +See the [CLI reference](../reference/cli.md#cache-shared-state) for the safety floors and +[`rpc_endpoint_tracker_gate_skips_total`](../reference/metrics.md#endpoint-scoped-rpc_endpoint_) +for how to confirm it is working. ## Observability diff --git a/docs/reference/cli.md b/docs/reference/cli.md index 5d5fdb2..0c4d65b 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -68,15 +68,23 @@ See [Failover & retry](../configuration/failover/index.md). ## Polling relief Each configured upstream gets its own chain tracker, which polls it for the latest -block independently of the relays you send. These two flags lower that load; both -are process-wide. Watch the effect on +block independently of the relays you send. These flags lower that load; all are +process-wide. Watch the effect on [`rpc_endpoint_tracker_requests_total`](metrics.md#endpoint-scoped-rpc_endpoint_), and confirm what is live via `HashPolling` in `GET /debug/endpoint-state`. | Flag | Default | Description | | --- | --- | --- | -| `--enable-fork-detection` | `false` | Turn on block-hash polling (reorg detection on upstreams). Off by default — the larger of the two savings. | +| `--enable-fork-detection` | `false` | Turn on block-hash polling (reorg detection on upstreams). Off by default — the larger of the per-pod savings. | | `--chain-tracker-poll-divisor` | `2` | The tracker polls every `avgBlockTime ÷ divisor`. `1` halves the polling rate. Allowed `[1,8]`; out-of-range reverts to the default. | +| `--shared-state` (with `--cache-be`) | `false` | Share poll observations across router replicas through the cache, so an upstream is polled about once per interval **fleet-wide** rather than once per pod. See [Cache & shared state](#cache-shared-state). | + +The tracker also skips a poll whenever something else has already kept the upstream's +tip fresh — the relays it served in the last block time, or (with `--shared-state`) a +poll another replica made. Skipped ticks show up on +[`rpc_endpoint_tracker_gate_skips_total`](metrics.md#endpoint-scoped-rpc_endpoint_) by +`source`, so `requests_total` falling while `gate_skips_total` rises is the relief +working, not the tracker stalling. ## Consistency tuning @@ -89,7 +97,17 @@ and confirm what is live via `HashPolling` in `GET /debug/endpoint-state`. | Flag | Default | Description | | --- | --- | --- | | `--cache-be` | — | Address of the cache server (e.g. `127.0.0.1:20100`). In Compose the cache address usually comes from `cache-be:` in the config instead. | -| `--shared-state` | `false` | Share consistency state across router instances via the cache (use with `--cache-be`). | +| `--shared-state` | `false` | Share state across router replicas through the cache (use with `--cache-be`): the consumer-consistency "seen block", **and** per-endpoint chain-tracker poll observations, so the fleet polls each upstream about once per interval instead of once per replica. | + +With `--shared-state`, every replica publishes each successful poll of an upstream to +the cache and, before polling, checks whether **another** replica polled that upstream +within the last block time. If one did, the tick is skipped and the peer's block is +adopted (it appears as `Source: peer` in `GET /debug/endpoint-state`). Three floors +keep it safe: a replica never borrows its own observation, so a single replica polls +exactly as without the flag; every replica still polls each upstream itself every few +ticks, so a broken path from one pod stays detectable; and an upstream that was +disabled is only re-enabled by that pod's own successful poll. Latency is never +shared — a peer's round-trip says nothing about this pod's path. ## WebSocket diff --git a/docs/reference/metrics.md b/docs/reference/metrics.md index 4587777..31abfcf 100644 --- a/docs/reference/metrics.md +++ b/docs/reference/metrics.md @@ -29,7 +29,7 @@ If you only graph a handful of things, graph these: | Are requests failing? | `smartrouter_total_errored`, `smartrouter_errors_total` (by `error_category` / `retryable`) | | How fast is it? | `smartrouter_end_to_end_latency_milliseconds` (histogram → p50/p95/p99) | | Is a node degrading? | `rpc_endpoint_overall_health`, `rpc_endpoint_end_to_end_latency_milliseconds`, `rpc_endpoint_latest_block` | -| How much load does the router itself put on my node? | `rpc_endpoint_tracker_requests_total` (by `kind`) — the router's own polling, separate from the relays you sent it | +| How much load does the router itself put on my node? | `rpc_endpoint_tracker_requests_total` (by `kind`) — the router's own polling, separate from the relays you sent it; `rpc_endpoint_tracker_gate_skips_total` (by `source`) — the polls it chose not to send | | Is failover working hard? | `smartrouter_retries_total`, `smartrouter_hedge_total` | | Is an upstream rate-limiting us? | `smartrouter_rate_limit_holdoffs_total` (by `provider` / `event`), `smartrouter_rate_limit_holdoff_seconds` | | Is the cache earning its keep? | `smartrouter_cache_success_total` / `smartrouter_cache_requests_total` | @@ -202,6 +202,13 @@ They split into **endpoint-scoped** (`rpc_endpoint_*`) and **router-scoped** | `rpc_endpoint_fetch_latest_success` | Counter | `spec`, `apiInterface`, `endpoint_id` | New-block **detections** by the chain tracker, not successful requests. | | `rpc_endpoint_fetch_block_success` | Counter | `spec`, `apiInterface`, `endpoint_id` | Successful specific-block fetches. | | `rpc_endpoint_tracker_requests_total` | Counter | `spec`, `apiInterface`, `endpoint_id`, `kind` | Requests the chain tracker actually sent upstream, by `kind` — `latest_block`, or `block_hash` (zero unless [`--enable-fork-detection`](cli.md#polling-relief)). The only metric that measures tracker **request volume**. | +| `rpc_endpoint_tracker_gate_skips_total` | Counter | `spec`, `apiInterface`, `endpoint_id`, `source` | Poll ticks the tracker skipped because the upstream's tip was already fresh, by `source` — `relay` (served traffic in the last block time) or `peer` (another replica's poll, shared via [`--shared-state`](cli.md#cache-shared-state)). The complement of `requests_total`: together they account for every tick. | + +On a multi-replica deployment with `--shared-state`, the fleet-wide +`sum(rate(rpc_endpoint_tracker_requests_total{kind="latest_block"}[5m]))` for an upstream +should settle near a single replica's worth, and `source="peer"` skips should be non-zero on +every replica. `peer` stuck at zero with the flag on means the replicas are not reaching the +same cache. Two caveats when reading the chain-tracker counters: From abc07e741741a187fce4ce3cecf683367c68e46b Mon Sep 17 00:00:00 2001 From: Nimrod Teich Date: Sun, 23 Aug 2026 13:41:39 +0300 Subject: [PATCH 2/2] docs(cli): --chain-tracker-poll-divisor allows a cadence slower than block time MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit smart-router#323 widened the range from [1,8] to [0.25,8]. At 0.25 the tracker polls once per four block times — 8x below the built-in cadence, and the relief that matters on a fast chain. The old floor of 1 was justified by the claim that below it the avgBlockTime-derived windows start to bind. They do not: the staleness window is max(10 x avgBlockTime, 2s) and does not move with this flag — the flag moves the other side of that comparison. Both common cases stay well inside it; the exposure is the seam where an upstream trips the traffic gate and then goes quiet, refreshed by neither polls nor relays. That seam is a property of the product of this flag and the gate's skip budget, so the router warns per chain at startup instead of refusing to start. Documented verbatim, with what to do about it, so the line is read as guidance rather than a failed boot. --- docs/reference/cli.md | 38 +++++++++++++++++++++++++++++++++++++- 1 file changed, 37 insertions(+), 1 deletion(-) diff --git a/docs/reference/cli.md b/docs/reference/cli.md index 0c4d65b..480407f 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -76,7 +76,7 @@ and confirm what is live via `HashPolling` in `GET /debug/endpoint-state`. | Flag | Default | Description | | --- | --- | --- | | `--enable-fork-detection` | `false` | Turn on block-hash polling (reorg detection on upstreams). Off by default — the larger of the per-pod savings. | -| `--chain-tracker-poll-divisor` | `2` | The tracker polls every `avgBlockTime ÷ divisor`. `1` halves the polling rate. Allowed `[1,8]`; out-of-range reverts to the default. | +| `--chain-tracker-poll-divisor` | `2` | The tracker polls every `avgBlockTime ÷ divisor`. `1` halves the polling rate; `0.25` — one poll per **four** block times — cuts it eightfold. Allowed `[0.25,8]`; out-of-range reverts to the default. See [Polling slower than the chain](#polling-slower-than-the-chain). | | `--shared-state` (with `--cache-be`) | `false` | Share poll observations across router replicas through the cache, so an upstream is polled about once per interval **fleet-wide** rather than once per pod. See [Cache & shared state](#cache-shared-state). | The tracker also skips a poll whenever something else has already kept the upstream's @@ -86,6 +86,42 @@ poll another replica made. Skipped ticks show up on `source`, so `requests_total` falling while `gate_skips_total` rises is the relief working, not the tracker stalling. +### Polling slower than the chain + +A divisor below `1` polls less often than the chain produces blocks, which is where +the relief is on a fast chain. Measured requests/min per endpoint: + +| Chain | `2` (default) | `1` | `0.5` | `0.25` | +| --- | --- | --- | --- | --- | +| Aptos | 600 | 300 | 150 | 75 | +| Solana | 300 | 150 | 75 | 37.5 | +| Base | 60 | 30 | 15 | 7.5 | +| Ethereum | 9.2 | 4.6 | 2.3 | 1.2 | + +What bounds the low end is the staleness window — `max(10 × avgBlockTime, 2s)`, past +which an observation stops counting for consensus, the tip reads unknown, and the probe +scores a healthy upstream not-alive. The window does **not** move with this flag; the +flag moves how long a tip can go unrefreshed, the other side of that comparison. Both +common cases stay well inside it: + +- **Idle upstream** — nothing but its own poll refreshes the tip, so the gap *is* the + interval: 4 × `avgBlockTime` at `0.25`, against a 10× window. +- **Served upstream** — relays refresh the same tip, so traffic bounds the gap. + +The seam is between them: an upstream that trips the traffic gate and *then* goes quiet +is refreshed by neither, and the worst-case gap becomes `(maxRelaySkips + 1) × interval` +— 20 × `avgBlockTime` at `0.25`. That is a property of the **product** of this flag and +the gate's skip budget, not of either alone, so the router warns once per chain at +startup rather than refusing to start: + +``` +WRN poll cadence can outrun the staleness window when the traffic gate skips + pollInterval=60s worstCaseGapBetweenPolls=5m0s stalenessWindow=2m30s maxRelaySkips=4 +``` + +It is a line to read, not a failure — the configuration is safe for both common cases. +If you see it on a chain with bursty traffic, step back toward `0.5` or `1`. + ## Consistency tuning | Flag | Default | Description |