Skip to content

feat(gateway): choose sites by published saturation per capacity - #289

Draft
hexfusion wants to merge 10 commits into
praxis-proxy:mainfrom
hexfusion:pr/gateway-site-selection
Draft

hexfusion wants to merge 10 commits into
praxis-proxy:mainfrom
hexfusion:pr/gateway-site-selection

Conversation

@hexfusion

Copy link
Copy Markdown
Collaborator

Summary

Stacked on #286. The gateway picks a site from grid_provider_saturation_ratio against its capacity, rather than ranking on a single queue-depth metric name, and sheds when no site has room. The operator keeps deriving the ratio: capacity is spec.maxRunning times ready endpoints, and maxRunning lives only in the peer site's Kubernetes API.

What changed

  • Selection: with three or more sites with room, draw two by capacity and keep the lower saturation. With two, draw weighted, clamped to 0.1 and 0.9. With one, take it.
  • No site with room: draw by capacity among the healthy sites, tied on the best polled queue depth.
  • Shed above a saturation of 1.05, route again below 0.95, decided per snapshot.
  • Load metric: accept llm_d_epp_average_queue_size as well as inference_pool_average_queue_size, the only name today's gateway reads.
  • Readiness exclusion and the decision counters, with an opt-in metrics listener and its chart support, carried from the earlier draft.
  • Demotion waits until praxis has reported cluster health, through a new health.observed().
  • Errors: a shed answers 429 rate_limit_exceeded with code capacity_exhausted, no healthy site stays 503 no_healthy_site, an unknown model is 404 model_not_found.

Testing

Gateway and root workspaces: fmt, clippy and tests pass. 3122 root and 98 gateway tests.

Partition simulation against the gateway binary, three sites at 128, 128 and 64 slots, 200ms service, both binaries built on praxis 0.7.3 so selection is the only difference.

clients build hub, 0.40 west, 0.40 east, 0.20 peak 10s window requests overshoot
64 before 1.00 0.00 0.00 1.00 14272 0
64 after 0.38 0.39 0.23 0.46 14272 0
128 before 1.00 0.00 0.00 1.00 28544 0
128 after 0.40 0.37 0.24 0.48 28508 49
256 before 1.00 0.00 0.00 1.00 28655 24517
256 after 0.34 0.46 0.20 0.48 42290 15199

No request failed in any run. Saturation alone tracks capacity share at all three loads, and converges as load rises, so the published latency series stays observe-only and no delay term is added.

The lab matches the before column. Across a sweep from 1 to 12 requests per second every Qwen3 request went to hq-east, which queued up to 193 while two sites sat idle. Three of the four sites published no load signal, because the operator's EPP scrape used a token bound to another pod. With nothing to rank, candidates tie and the gateway takes the front of a local-first order. That scrape is fixed separately, but the deployed gateway reads only inference_pool_average_queue_size and the lab's router 0.11.0 EPP exports llm_d_epp_average_queue_size, so it would tie again.

Limitations

  • Under capacity the throughput is the same. At 64 and 128 clients the clients are the limit, not the grid, so both builds serve the same requests. The 48 percent gain appears only at 256, where the old path caps at one site's 128 slots.
  • At 128 clients this takes 49 overshoots where the old path takes none. Pinning 128 clients onto a 128-slot site is an exact fit by accident, and spreading puts a few requests on the 64-slot site past full. It is 0.17 percent of the run and nothing failed.
  • Overshoot falls 38 percent at 256 clients but does not go away. Selection acts on a ratio up to one poll stale, so it cannot promise a site has a free slot on arrival. Removing overshoot needs admission control, not better selection.
  • The two equal sites split 0.34 and 0.46 at 256 clients. Inside the bar the row sets, but the weighted draw is not tight between equals under heavy load.
  • Saturation carries no delay term, and no KV or EPP-gate term.
  • The overflow draw is unbounded.
  • A signal is treated as fresh for two polls, not the wider bound the design derives from poll, scrape, poll timeout and one more poll.

The scraper token was bound to the operator pod, which runs as the
grid-operator ServiceAccount, and the API server binds a token only to
a pod running as the token's own ServiceAccount. Every TokenRequest
failed with 400, so no serviceAccountToken scrape ever ran, and the
failure logged only at debug.

Mint it unbound, keeping the 10-minute lifetime, the API server's
default audience the EPP's TokenReview checks, and reuse at two thirds.
A failed mint now logs at WARN when the error changes, and at INFO when
minting recovers. The chart drops the pod name and uid env it read.

The same mint code is in the base of praxis-proxy#286, which rebases on this.

Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
The operator decides from each scrape of a provider's EPP whether it can serve
now, publishes the verdict as grid_provider_ready, writes it to the provider's
Ready condition on change, and shows it as one STATUS column, like nodes.

Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
…nd saturation

grid_provider_in_flight_requests is the larger of the EPP's per-endpoint count
and the pool averages times ready endpoints, plus flow control's queue.
spec.maxRunning now counts one endpoint; grid_provider_capacity_requests is it
times fresh ready endpoints, and grid_provider_saturation_ratio is in-flight over
capacity for a finite non-negative count. The grid-operator chart passes
grid.providers.<name>.maxRunning through. Series take the signal contract names.
Reinstall with per-endpoint maxRunning.

Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
TTFT p50 and p90, TPOT, prefill seconds per token, and error ratio over the last
30s, from deltas of the EPP's request histograms, once 20 requests complete.

Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
…ub keeps from peers

The provider series are gauges on /metrics, labeled grid_site and grid_provider,
this site's and its peers', each in its own family. A hub keeps from a peer only
the contract names and the EPP pool averages, DNS-1123 grid_provider values,
finite non-negative values, and 64 providers; refusals count in
grid_peer_signals_refused_total{peer,reason}.

Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
Ready uses the status contract's reasons: NoEndpointsReady, NoLivenessCheck
for a scrape without the pool's ready-endpoint series, ScrapeTimedOut,
ScrapeUnauthorized, TLSHandshakeFailed or ScrapeFailed with the class in the
message, and AwaitingFirstScrape (Unknown). A change logs WARN when the
provider turns not ready. Scrapes count in grid_provider_scrape_total and
grid_provider_last_scrape_success_timestamp_seconds. A not-ready provider keeps
a ready=0 row. The mock provider exposes a ready count.

Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
…ng sites

A provider with an empty site selector is placed at its own site only. The
serving config refuses candidates whose site is not a DNS-1123 label, warning
per GridNetwork on change, and derives the gateway's load window from the
scrape, poll, and poll timeout plus one poll of slack, 17s at a 5s scrape.

Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
…Ready

The EPP's ready-endpoint count is endpoints with fresh metrics, so a saturated
engine whose /metrics answers late reads as zero while it serves, and excluding
it herds its load onto the other sites. Zero counts now exclude only when the
EPP also recorded no first token, completion or usage report for the pool in
the last 30s. The running averages are no evidence: they freeze when the count
is zero.

Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
The provider scrape went through the peer-relay parser, which drops counters
and histogram parts, so no latency, error ratio or engine-answer evidence was
ever computed. A local scrape now keeps every sample for this site's windows
and republishes gauges only, as before.

Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
Each site's load is rho, the grid_provider_saturation_ratio its operator
publishes, read when the snapshot is built. Among healthy sites with room (rho
below 1): two choices drawn by capacity at three or more, taking the lower rho; a
draw weighted by capacity over 1 + rho at two, each share kept within 0.1 to 0.9;
or the only one. With none, a capacity draw among the healthy sites tied on the
best polled queue depth, which also covers sites that publish no load. Capacity
is unknown, not 1, when unpublished.

Selection no longer takes the front match of a tied list, which pinned every
request to the first local cluster whenever sites tied, at zero load or at the
same queue depth.

A candidate whose latest grid_provider_ready sample is 0 is excluded, and one on
a cluster praxis reports with no healthy endpoint is ordered last. Demotion
applies only once praxis has reported health, so an unobserved registry demotes
nothing.

A model sheds when every healthy site serving it holds rho of at least 1.05, and
routes again once one reaches 0.95, decided per snapshot. A shed answers 429 with
Retry-After and an OpenAI-style error, type rate_limit_exceeded and code
capacity_exhausted, so a client backs off as it does for any overload. No healthy
site answers 503 no_healthy_site, an outage rather than load, and an unknown
model 404 model_not_found.

grid_route_decisions_total counts every outcome by site and a closed reason set,
served with each site's score on an opt-in TLS metrics listener that the chart
exposes with a Service and ServiceMonitor.

A Narrow stage can clear sites with room before the draw, never add one. It sees
each site's rho and the recent latency its operator publishes, which ranking does
not use. Draws mix a per-request turn that starts at a random seed, so replicas
do not draw in step. One request weighs at most 64 sites with room.

docs/site-selection.md is a tuning guide for this path.

Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
@coderabbitai

coderabbitai Bot commented Oct 5, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant