Skip to content

fix(slo): format the budget percent from $values, and stop paging on eval errors - #15

Merged
chris13524 merged 1 commit into
mainfrom
fix/slo-alert-value-formatting
Oct 1, 2026
Merged

chris13524 merged 1 commit into
mainfrom
fix/slo-alert-value-formatting

Conversation

@chris13524

Copy link
Copy Markdown
Member

What happened

A P3 reached blockchain-api's on-call (twice, a day apart) reading:

[no value] burned %!f(string=)% of its 30-day error budget in the last 3d, and is still burning as of the last 6h. Tier limit is 10% per 3d — 1x the sustainable pace, which exhausts the month in 30 days. Objective: [no value].

Grafana's alertname on it was DatasourceError. Two independent defects in this library are visible in that one line.

$value is a string, not a float

slo/rule.libsonnet formatted the budget percentage with {{ printf "%.1f" $value }}. In Grafana alerting $value is a string — the human-readable rendering of every captured value, [ var='B' labels={...} value=12.3 ] — so a float verb applied to it produces %!f(string=...). The percent has therefore never rendered since the annotations were rewritten in e8621e6; the empty (string=) in the page above is the same bug with nothing captured.

The float lives at $values.<refId>.Value. B is the reduce node alert_rule.libsonnet builds between the query (A) and the threshold (C), so it carries the expression's value. Wrapped in with so an instance that captured nothing renders ? instead of a fresh class of garbage.

exec_err_state was left at Error

The rule set no_data_state = OK but never passed exec_err_state, inheriting the library default of Error (alert_rule.libsonnet:83). A burn-rate rule measures a 30-day budget; one failed evaluation is not evidence about that budget. The default is what turned a query failure into a page — Grafana raises a DatasourceError alert carrying this rule's annotations, and a DatasourceError instance has none of the query's labels, which is where both [no value]s come from.

Now OK. The cost is that a persistently broken rule goes quiet rather than shouting, which is acceptable only because rule health is a different signal for a different audience — it belongs on grafana_alerting_rule_evaluation_failures_total, reaching whoever owns the monitoring stack rather than whoever is on call for the service. That alert does not exist yet; worth a follow-up.

Tests

Both properties are pinned in slo/tests/smoke.jsonnet, the only test here that renders the Grafana rule at all — promtool evaluates PromQL and can see neither an annotation template nor a failure state. The $value guard counts occurrences rather than matching the prose, so it survives rewording and catches any float verb, not just %.1f.

Each guard was mutation-checked by reintroducing the defect it names; both fail with the intended message and exit 1.

Full CI suite run locally on go-jsonnet 0.22.0 and promtool 3.1.0 (the versions ci.yml pins): smoke renders, all three promtool suites SUCCESS.

Rendered output for the fast tier:

Prod - SLO fast burn: {{ $labels.sli }}{{ if $labels.route }} on route {{ $labels.route }}{{ end }} — {{ with $values.B }}{{ printf "%.1f" .Value }}{{ else }}?{{ end }}% of 30-day budget in 1h (limit 2%)

Follow-ups, not in this PR

  • blockchain-api needs a submodule bump to pick this up. Note its main currently pins 7d57d4d, which is not on this repo's main — it only exists on feat/slo-decorate-hook. That wants untangling as part of the bump; this PR targets main, where the bug was introduced.
  • pay-core consumes the same library but its checked-in consumer still uses the older, correct {{ $value }} wording, so it is unaffected by the first defect and fixed for free on the second.
  • Root cause of the datasource error itself is still open. This PR stops the garbled page and stops the page; it does not explain why the query errored. The rule's Health/Error string in Grafana will say whether it fails every evaluation (stale datasource UID) or intermittently (the 30d expression hitting a query timeout or max_samples).

🤖 Generated with Claude Code

…eval errors

A P3 reached blockchain-api's on-call reading:

  [no value] burned %!f(string=)% of its 30-day error budget in the last 3d
  ... Objective: [no value].

Two independent defects, both visible in that one line.

`$value` is a STRING in Grafana alerting — the rendering of every captured
value, `[ var='B' labels={...} value=12.3 ]` — not a float. `printf "%.1f"`
applied to it yields `%!f(string=...)`, so the percent has never rendered since
the annotations were rewritten in e8621e6; the empty `(string=)` is the same bug
with nothing captured. The float is `$values.B.Value`, `B` being the reduce node
alert_rule.libsonnet puts between the query and the threshold. Guarded with
`with`, so an instance that captured nothing says `?` rather than more garbage.

`exec_err_state` was left at the library default of `Error`. A burn-rate rule
measures a 30-day budget and a single failed evaluation is not evidence about
it, but the default makes Grafana raise a DatasourceError carrying these
annotations — which is how a query failure came to page a service on-call with a
budget message whose every templated field was empty. The `[no value]`s are that:
a DatasourceError instance has none of the query's labels. Set to `OK`; rule
health belongs on grafana_alerting_rule_evaluation_failures_total, which reaches
the monitoring owners instead of the service's on-call.

Both are pinned in smoke.jsonnet, the only test here that renders the Grafana
rule at all — promtool covers the PromQL and can see neither. Each guard was
mutation-checked by reintroducing the defect it names.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@chris13524

Copy link
Copy Markdown
Member Author

[AI] Root cause of the underlying datasource error is now confirmed from Grafana's alert state history. It is not a query timeout or max_samples, which is what I speculated when opening this — it is a 414 Request-URI Too Large.

[sse.dataQueryError] failed to execute query [A]: unexpected response with status code 414:
<html><head><title>414 Request-URI Too Large</title></head>...

The prod-blockchain-api-amp datasource (7BAO6wDSk) is configured httpMethod: "GET", so Grafana puts the whole PromQL expression in the URL. The budget tier's rendered expression is 17,648 characters (slo_rules_all.generated.json), which AMP rejects.

The other two tiers pin the threshold neatly: slo-burn-rate-fast is 11,912 chars and slo-burn-rate-slow is 11,917, and both evaluate fine — the slow tier is currently Alerting on Monad Mainnet with real values. Only the 3d tier crosses the limit. That is also why only the P3 ever misfired.

The 3d rule is in permanent Error, not intermittent: every terraform apply resets it to Normal (Updated), the next evaluation errors, and it walks Pending (Error) → Error and re-notifies. It has never once evaluated since it was provisioned.

This raises the stakes on the exec_err_state = OK half of this PR. The "persistently broken rule goes quiet" cost I described as a trade-off is the actual present state of this rule — merging this alone would silence a 3d SLO tier that has never worked, with nothing to say so. Suggested landing order:

  1. httpMethod = "POST" on the AMP datasource in blockchain-api (terraform/monitoring/data_sources.tf, both the AMG resource and the cloud_amp twin) — this is what actually makes the rule evaluate.
  2. The rule-health alert on grafana_alerting_rule_evaluation_failures_total.
  3. This PR + the submodule bump.

Nothing here changes the diff — both defects are real and the fix stands on its own. Flagging the sequencing.

@chris13524
chris13524 merged commit 9b6555f into main Oct 1, 2026
3 checks passed
@chris13524
chris13524 deleted the fix/slo-alert-value-formatting branch October 1, 2026 13:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants