fix(slo): format the budget percent from $values, and stop paging on eval errors - #15
Conversation
…eval errors
A P3 reached blockchain-api's on-call reading:
[no value] burned %!f(string=)% of its 30-day error budget in the last 3d
... Objective: [no value].
Two independent defects, both visible in that one line.
`$value` is a STRING in Grafana alerting — the rendering of every captured
value, `[ var='B' labels={...} value=12.3 ]` — not a float. `printf "%.1f"`
applied to it yields `%!f(string=...)`, so the percent has never rendered since
the annotations were rewritten in e8621e6; the empty `(string=)` is the same bug
with nothing captured. The float is `$values.B.Value`, `B` being the reduce node
alert_rule.libsonnet puts between the query and the threshold. Guarded with
`with`, so an instance that captured nothing says `?` rather than more garbage.
`exec_err_state` was left at the library default of `Error`. A burn-rate rule
measures a 30-day budget and a single failed evaluation is not evidence about
it, but the default makes Grafana raise a DatasourceError carrying these
annotations — which is how a query failure came to page a service on-call with a
budget message whose every templated field was empty. The `[no value]`s are that:
a DatasourceError instance has none of the query's labels. Set to `OK`; rule
health belongs on grafana_alerting_rule_evaluation_failures_total, which reaches
the monitoring owners instead of the service's on-call.
Both are pinned in smoke.jsonnet, the only test here that renders the Grafana
rule at all — promtool covers the PromQL and can see neither. Each guard was
mutation-checked by reintroducing the defect it names.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
[AI] Root cause of the underlying datasource error is now confirmed from Grafana's alert state history. It is not a query timeout or The The other two tiers pin the threshold neatly: The 3d rule is in permanent Error, not intermittent: every terraform apply resets it to This raises the stakes on the
Nothing here changes the diff — both defects are real and the fix stands on its own. Flagging the sequencing. |
What happened
A P3 reached blockchain-api's on-call (twice, a day apart) reading:
Grafana's
alertnameon it wasDatasourceError. Two independent defects in this library are visible in that one line.$valueis a string, not a floatslo/rule.libsonnetformatted the budget percentage with{{ printf "%.1f" $value }}. In Grafana alerting$valueis a string — the human-readable rendering of every captured value,[ var='B' labels={...} value=12.3 ]— so a float verb applied to it produces%!f(string=...). The percent has therefore never rendered since the annotations were rewritten in e8621e6; the empty(string=)in the page above is the same bug with nothing captured.The float lives at
$values.<refId>.Value.Bis the reduce nodealert_rule.libsonnetbuilds between the query (A) and the threshold (C), so it carries the expression's value. Wrapped inwithso an instance that captured nothing renders?instead of a fresh class of garbage.exec_err_statewas left atErrorThe rule set
no_data_state = OKbut never passedexec_err_state, inheriting the library default ofError(alert_rule.libsonnet:83). A burn-rate rule measures a 30-day budget; one failed evaluation is not evidence about that budget. The default is what turned a query failure into a page — Grafana raises aDatasourceErroralert carrying this rule's annotations, and aDatasourceErrorinstance has none of the query's labels, which is where both[no value]s come from.Now
OK. The cost is that a persistently broken rule goes quiet rather than shouting, which is acceptable only because rule health is a different signal for a different audience — it belongs ongrafana_alerting_rule_evaluation_failures_total, reaching whoever owns the monitoring stack rather than whoever is on call for the service. That alert does not exist yet; worth a follow-up.Tests
Both properties are pinned in
slo/tests/smoke.jsonnet, the only test here that renders the Grafana rule at all — promtool evaluates PromQL and can see neither an annotation template nor a failure state. The$valueguard counts occurrences rather than matching the prose, so it survives rewording and catches any float verb, not just%.1f.Each guard was mutation-checked by reintroducing the defect it names; both fail with the intended message and exit 1.
Full CI suite run locally on go-jsonnet 0.22.0 and promtool 3.1.0 (the versions
ci.ymlpins): smoke renders, all three promtool suites SUCCESS.Rendered output for the fast tier:
Follow-ups, not in this PR
maincurrently pins7d57d4d, which is not on this repo'smain— it only exists onfeat/slo-decorate-hook. That wants untangling as part of the bump; this PR targetsmain, where the bug was introduced.{{ $value }}wording, so it is unaffected by the first defect and fixed for free on the second.max_samples).🤖 Generated with Claude Code