Skip to content

feat(operator)!: publish provider readiness, load, capacity, and latency - #286

Draft
hexfusion wants to merge 9 commits into
praxis-proxy:mainfrom
hexfusion:pr/operator-provider-signals
Draft

hexfusion wants to merge 9 commits into
praxis-proxy:mainfrom
hexfusion:pr/operator-provider-signals

Conversation

@hexfusion

@hexfusion hexfusion commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

The operator decides whether each provider can serve now and says why, as the provider's Ready condition and as grid_provider_ready. It also publishes how many requests each provider holds, its capacity, their ratio, and recent latency, on its signals endpoint and on /metrics. Every input is the EPP's /metrics. Stacked on #288.

What changed

  • Readiness: Ready uses the status contract's reasons. NoEndpointsReady after two zero counts with no engine answer in 30s, since the EPP counts endpoints with fresh metrics and a saturated engine reads as zero, NoLivenessCheck when a scrape answers without the pool's ready-endpoint series, and ScrapeTimedOut, ScrapeUnauthorized, TLSHandshakeFailed or ScrapeFailed (class in the message) once nothing succeeds in the window. AwaitingFirstScrape (Unknown) before the first success. A change logs WARN when the provider turns not ready, INFO otherwise.
  • Scrapes: grid_provider_scrape_total{grid_provider,result} and grid_provider_last_scrape_success_timestamp_seconds{grid_provider}.
  • In-flight: grid_provider_in_flight_requests, the larger of the EPP's per-endpoint count and the pool averages times ready endpoints, plus flow control's queue.
  • Capacity, breaking: spec.maxRunning counts one endpoint. grid_provider_capacity_requests is it times fresh ready endpoints. The grid-operator chart now passes grid.providers..maxRunning through.
  • Saturation: grid_provider_saturation_ratio, in-flight over capacity, only for a finite non-negative count.
  • Latency: TTFT p50 and p90, TPOT, prefill seconds per token, and error ratio over 30s, from the EPP's histograms, once 20 requests complete.
  • /metrics: the provider series as gauges labeled grid_site and grid_provider, this site's and its peers'.
  • Peer bounds: a hub keeps from a peer only the contract names and the EPP pool averages, DNS-1123 grid_provider values, finite non-negative values, and 64 providers. Refusals count in grid_peer_signals_refused_total{peer,reason}. A peer's custom signalNames are dropped.
  • Overlay: a provider with an empty site selector is placed at its own site only.
  • Serving config: candidates whose site is not a DNS-1123 label are refused, warned per GridNetwork on change.
  • Load window: derived from the scrape, poll and poll timeout plus one poll of slack, 17s at a 5s scrape, instead of 30s.
  • Mock provider: exposes a ready count, as an EPP does.

Testing

  • Operator: clippy, fmt and tests, 1576 passed.
  • Mock provider tests, 59 passed.
  • Helm: lint --strict on every chart, unittest, and verify-helm-chart, 349 checks.

Limitations

  • No e2e topology scrapes an EPP, so readiness, in-flight, capacity and latency are covered by unit tests only.
  • The scraper token now reaches every provider with a metricsConfig. Narrowing its audience and targets is a separate change.
  • A mock provider image older than this change reads NoLivenessCheck.
  • Reinstall with per-endpoint maxRunning.

@coderabbitai

coderabbitai Bot commented Oct 4, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

The scraper token was bound to the operator pod, which runs as the
grid-operator ServiceAccount, and the API server binds a token only to
a pod running as the token's own ServiceAccount. Every TokenRequest
failed with 400, so no serviceAccountToken scrape ever ran, and the
failure logged only at debug.

Mint it unbound, keeping the 10-minute lifetime, the API server's
default audience the EPP's TokenReview checks, and reuse at two thirds.
A failed mint now logs at WARN when the error changes, and at INFO when
minting recovers. The chart drops the pod name and uid env it read.

The same mint code is in the base of praxis-proxy#286, which rebases on this.

Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
@hexfusion
hexfusion force-pushed the pr/operator-provider-signals branch from 2859e64 to e429fdb Compare October 4, 2026 22:00
The operator decides from each scrape of a provider's EPP whether it can serve
now, publishes the verdict as grid_provider_ready, writes it to the provider's
Ready condition on change, and shows it as one STATUS column, like nodes.

Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
…nd saturation

grid_provider_in_flight_requests is the larger of the EPP's per-endpoint count
and the pool averages times ready endpoints, plus flow control's queue.
spec.maxRunning now counts one endpoint; grid_provider_capacity_requests is it
times fresh ready endpoints, and grid_provider_saturation_ratio is in-flight over
capacity for a finite non-negative count. The grid-operator chart passes
grid.providers.<name>.maxRunning through. Series take the signal contract names.
Reinstall with per-endpoint maxRunning.

Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
TTFT p50 and p90, TPOT, prefill seconds per token, and error ratio over the last
30s, from deltas of the EPP's request histograms, once 20 requests complete.

Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
…ub keeps from peers

The provider series are gauges on /metrics, labeled grid_site and grid_provider,
this site's and its peers', each in its own family. A hub keeps from a peer only
the contract names and the EPP pool averages, DNS-1123 grid_provider values,
finite non-negative values, and 64 providers; refusals count in
grid_peer_signals_refused_total{peer,reason}.

Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
Ready uses the status contract's reasons: NoEndpointsReady, NoLivenessCheck
for a scrape without the pool's ready-endpoint series, ScrapeTimedOut,
ScrapeUnauthorized, TLSHandshakeFailed or ScrapeFailed with the class in the
message, and AwaitingFirstScrape (Unknown). A change logs WARN when the
provider turns not ready. Scrapes count in grid_provider_scrape_total and
grid_provider_last_scrape_success_timestamp_seconds. A not-ready provider keeps
a ready=0 row. The mock provider exposes a ready count.

Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
…ng sites

A provider with an empty site selector is placed at its own site only. The
serving config refuses candidates whose site is not a DNS-1123 label, warning
per GridNetwork on change, and derives the gateway's load window from the
scrape, poll, and poll timeout plus one poll of slack, 17s at a 5s scrape.

Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
@hexfusion
hexfusion force-pushed the pr/operator-provider-signals branch from e429fdb to 2802be4 Compare October 4, 2026 22:25
…Ready

The EPP's ready-endpoint count is endpoints with fresh metrics, so a saturated
engine whose /metrics answers late reads as zero while it serves, and excluding
it herds its load onto the other sites. Zero counts now exclude only when the
EPP also recorded no first token, completion or usage report for the pool in
the last 30s. The running averages are no evidence: they freeze when the count
is zero.

Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
@hexfusion
hexfusion force-pushed the pr/operator-provider-signals branch from 1eca09f to e422e0c Compare October 4, 2026 23:10
The provider scrape went through the peer-relay parser, which drops counters
and histogram parts, so no latency, error ratio or engine-answer evidence was
ever computed. A local scrape now keeps every sample for this site's windows
and republishes gauges only, as before.

Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant