Conversation
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Comment |
The scraper token was bound to the operator pod, which runs as the grid-operator ServiceAccount, and the API server binds a token only to a pod running as the token's own ServiceAccount. Every TokenRequest failed with 400, so no serviceAccountToken scrape ever ran, and the failure logged only at debug. Mint it unbound, keeping the 10-minute lifetime, the API server's default audience the EPP's TokenReview checks, and reuse at two thirds. A failed mint now logs at WARN when the error changes, and at INFO when minting recovers. The chart drops the pod name and uid env it read. The same mint code is in the base of praxis-proxy#286, which rebases on this. Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
hexfusion
force-pushed
the
pr/operator-provider-signals
branch
from
October 4, 2026 22:00
2859e64 to
e429fdb
Compare
The operator decides from each scrape of a provider's EPP whether it can serve now, publishes the verdict as grid_provider_ready, writes it to the provider's Ready condition on change, and shows it as one STATUS column, like nodes. Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
…nd saturation grid_provider_in_flight_requests is the larger of the EPP's per-endpoint count and the pool averages times ready endpoints, plus flow control's queue. spec.maxRunning now counts one endpoint; grid_provider_capacity_requests is it times fresh ready endpoints, and grid_provider_saturation_ratio is in-flight over capacity for a finite non-negative count. The grid-operator chart passes grid.providers.<name>.maxRunning through. Series take the signal contract names. Reinstall with per-endpoint maxRunning. Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
TTFT p50 and p90, TPOT, prefill seconds per token, and error ratio over the last 30s, from deltas of the EPP's request histograms, once 20 requests complete. Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
…ub keeps from peers
The provider series are gauges on /metrics, labeled grid_site and grid_provider,
this site's and its peers', each in its own family. A hub keeps from a peer only
the contract names and the EPP pool averages, DNS-1123 grid_provider values,
finite non-negative values, and 64 providers; refusals count in
grid_peer_signals_refused_total{peer,reason}.
Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
Ready uses the status contract's reasons: NoEndpointsReady, NoLivenessCheck for a scrape without the pool's ready-endpoint series, ScrapeTimedOut, ScrapeUnauthorized, TLSHandshakeFailed or ScrapeFailed with the class in the message, and AwaitingFirstScrape (Unknown). A change logs WARN when the provider turns not ready. Scrapes count in grid_provider_scrape_total and grid_provider_last_scrape_success_timestamp_seconds. A not-ready provider keeps a ready=0 row. The mock provider exposes a ready count. Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
…ng sites A provider with an empty site selector is placed at its own site only. The serving config refuses candidates whose site is not a DNS-1123 label, warning per GridNetwork on change, and derives the gateway's load window from the scrape, poll, and poll timeout plus one poll of slack, 17s at a 5s scrape. Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
hexfusion
force-pushed
the
pr/operator-provider-signals
branch
from
October 4, 2026 22:25
e429fdb to
2802be4
Compare
…Ready The EPP's ready-endpoint count is endpoints with fresh metrics, so a saturated engine whose /metrics answers late reads as zero while it serves, and excluding it herds its load onto the other sites. Zero counts now exclude only when the EPP also recorded no first token, completion or usage report for the pool in the last 30s. The running averages are no evidence: they freeze when the count is zero. Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
hexfusion
force-pushed
the
pr/operator-provider-signals
branch
from
October 4, 2026 23:10
1eca09f to
e422e0c
Compare
The provider scrape went through the peer-relay parser, which drops counters and histogram parts, so no latency, error ratio or engine-answer evidence was ever computed. A local scrape now keeps every sample for this site's windows and republishes gauges only, as before. Signed-off-by: Sam Batschelet <sbatsche@redhat.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The operator decides whether each provider can serve now and says why, as the provider's Ready condition and as grid_provider_ready. It also publishes how many requests each provider holds, its capacity, their ratio, and recent latency, on its signals endpoint and on /metrics. Every input is the EPP's /metrics. Stacked on #288.
What changed
Testing
Limitations