Conversation
The scraper token was bound to the operator pod, which runs as the grid-operator ServiceAccount, and the API server binds a token only to a pod running as the token's own ServiceAccount. Every TokenRequest failed with 400, so no serviceAccountToken scrape ever ran, and the failure logged only at debug. Mint it unbound, keeping the 10-minute lifetime, the API server's default audience the EPP's TokenReview checks, and reuse at two thirds. A failed mint now logs at WARN when the error changes, and at INFO when minting recovers. The chart drops the pod name and uid env it read. The same mint code is in the base of praxis-proxy#286, which rebases on this. Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
The operator decides from each scrape of a provider's EPP whether it can serve now, publishes the verdict as grid_provider_ready, writes it to the provider's Ready condition on change, and shows it as one STATUS column, like nodes. Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
…nd saturation grid_provider_in_flight_requests is the larger of the EPP's per-endpoint count and the pool averages times ready endpoints, plus flow control's queue. spec.maxRunning now counts one endpoint; grid_provider_capacity_requests is it times fresh ready endpoints, and grid_provider_saturation_ratio is in-flight over capacity for a finite non-negative count. The grid-operator chart passes grid.providers.<name>.maxRunning through. Series take the signal contract names. Reinstall with per-endpoint maxRunning. Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
TTFT p50 and p90, TPOT, prefill seconds per token, and error ratio over the last 30s, from deltas of the EPP's request histograms, once 20 requests complete. Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
…ub keeps from peers
The provider series are gauges on /metrics, labeled grid_site and grid_provider,
this site's and its peers', each in its own family. A hub keeps from a peer only
the contract names and the EPP pool averages, DNS-1123 grid_provider values,
finite non-negative values, and 64 providers; refusals count in
grid_peer_signals_refused_total{peer,reason}.
Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
Ready uses the status contract's reasons: NoEndpointsReady, NoLivenessCheck for a scrape without the pool's ready-endpoint series, ScrapeTimedOut, ScrapeUnauthorized, TLSHandshakeFailed or ScrapeFailed with the class in the message, and AwaitingFirstScrape (Unknown). A change logs WARN when the provider turns not ready. Scrapes count in grid_provider_scrape_total and grid_provider_last_scrape_success_timestamp_seconds. A not-ready provider keeps a ready=0 row. The mock provider exposes a ready count. Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
…ng sites A provider with an empty site selector is placed at its own site only. The serving config refuses candidates whose site is not a DNS-1123 label, warning per GridNetwork on change, and derives the gateway's load window from the scrape, poll, and poll timeout plus one poll of slack, 17s at a 5s scrape. Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
…Ready The EPP's ready-endpoint count is endpoints with fresh metrics, so a saturated engine whose /metrics answers late reads as zero while it serves, and excluding it herds its load onto the other sites. Zero counts now exclude only when the EPP also recorded no first token, completion or usage report for the pool in the last 30s. The running averages are no evidence: they freeze when the count is zero. Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
The provider scrape went through the peer-relay parser, which drops counters and histogram parts, so no latency, error ratio or engine-answer evidence was ever computed. A local scrape now keeps every sample for this site's windows and republishes gauges only, as before. Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
Each site's load is rho, the grid_provider_saturation_ratio its operator publishes, read when the snapshot is built. Among healthy sites with room (rho below 1): two choices drawn by capacity at three or more, taking the lower rho; a draw weighted by capacity over 1 + rho at two, each share kept within 0.1 to 0.9; or the only one. With none, a capacity draw among the healthy sites tied on the best polled queue depth, which also covers sites that publish no load. Capacity is unknown, not 1, when unpublished. Selection no longer takes the front match of a tied list, which pinned every request to the first local cluster whenever sites tied, at zero load or at the same queue depth. A candidate whose latest grid_provider_ready sample is 0 is excluded, and one on a cluster praxis reports with no healthy endpoint is ordered last. Demotion applies only once praxis has reported health, so an unobserved registry demotes nothing. A model sheds when every healthy site serving it holds rho of at least 1.05, and routes again once one reaches 0.95, decided per snapshot. A shed answers 429 with Retry-After and an OpenAI-style error, type rate_limit_exceeded and code capacity_exhausted, so a client backs off as it does for any overload. No healthy site answers 503 no_healthy_site, an outage rather than load, and an unknown model 404 model_not_found. grid_route_decisions_total counts every outcome by site and a closed reason set, served with each site's score on an opt-in TLS metrics listener that the chart exposes with a Service and ServiceMonitor. A Narrow stage can clear sites with room before the draw, never add one. It sees each site's rho and the recent latency its operator publishes, which ranking does not use. Draws mix a per-request turn that starts at a random seed, so replicas do not draw in step. One request weighs at most 64 sites with room. docs/site-selection.md is a tuning guide for this path. Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Stacked on #286. The gateway picks a site from grid_provider_saturation_ratio against its capacity, rather than ranking on a single queue-depth metric name, and sheds when no site has room. The operator keeps deriving the ratio: capacity is spec.maxRunning times ready endpoints, and maxRunning lives only in the peer site's Kubernetes API.
What changed
Testing
Gateway and root workspaces: fmt, clippy and tests pass. 3122 root and 98 gateway tests.
Partition simulation against the gateway binary, three sites at 128, 128 and 64 slots, 200ms service, both binaries built on praxis 0.7.3 so selection is the only difference.
No request failed in any run. Saturation alone tracks capacity share at all three loads, and converges as load rises, so the published latency series stays observe-only and no delay term is added.
The lab matches the before column. Across a sweep from 1 to 12 requests per second every Qwen3 request went to hq-east, which queued up to 193 while two sites sat idle. Three of the four sites published no load signal, because the operator's EPP scrape used a token bound to another pod. With nothing to rank, candidates tie and the gateway takes the front of a local-first order. That scrape is fixed separately, but the deployed gateway reads only inference_pool_average_queue_size and the lab's router 0.11.0 EPP exports llm_d_epp_average_queue_size, so it would tie again.
Limitations