Skip to content

feat(slo): reusable multiwindow multi-burn-rate SLO alerting - #12

Merged
chris13524 merged 1 commit into
mainfrom
feat/slo-burn-rate
Sep 16, 2026
Merged

chris13524 merged 1 commit into
mainfrom
feat/slo-burn-rate

Conversation

@chris13524

Copy link
Copy Markdown
Member

Extracted from pay-core and blockchain-api, which had each grown a copy of the same
burn-rate implementation. This is the shared home so they stop diverging — I created that
divergence last week by copying the file, so this closes it.

What's here

File
slo/burn_rate.libsonnet tier table, budget sizing, the PromQL conditions
slo/rule.libsonnet tiers + SLIs → Grafana Unified Alerting rule groups
slo/panels.libsonnet burn-rate chart + error-budget bar gauge
slo/tests/ promtool suites over the shipping expressions
instant_query.libsonnet make an alert rule's query instant rather than range
.github/workflows/ci.yml this repo's first CI

Why so little is parameterised

I diffed the two copies before extracting. Ignoring comments and whitespace, they
differed in two things: the label carrying the priority (priority vs
og_priority) and two sentences of annotation prose. So those are the only opts; the
rest was already identical. Every knob is a hidden field, so a consumer retunes with
burnRate + { tiers:: [...] } instead of forking.

Two bugs the extraction surfaced

Writing fixtures that didn't share the consumers' incidental metric shapes broke the
code immediately — which is the argument for a library having its own tests.

  1. events() depended on a label-set accident. Both copies selected the good and bad
    counters with a single __name__ matcher. increase() drops __name__, so the two
    collapse to the same label set and Prometheus rejects the vector — unless some other
    label distinguishes them. In blockchain-api one metric happens to carry
    provider_kind. Remove it and every SLO rule goes to exec_err_state together, from
    an unrelated metrics refactor. Fixed with the tag-and-union idiom max_of already
    uses; the fixtures deliberately give good/bad identical label sets so it cannot
    come back. (Fixed in blockchain-api separately.)

  2. sum(good) + sum(bad) — the form I tried first — yields nothing when either
    side has no series, which is exactly a dimension in total outage.

Testing

CI installs go-jsonnet (matching Terraform's alxrem/jsonnet provider, which embeds it)
and promtool, renders rules and both panels from fixtures, then evaluates 3 suites
(~4s). The suites test mechanics — window arithmetic, budget sizing, the noise floor,
expression splicing, the sparse-denominator hazard, panel/rule agreement. Every case is
mutation-checked.

Not a breaking change: consumers pin a submodule SHA, so nothing moves until they
bump it. Follow-ups will migrate blockchain-api and pay-core onto this and delete
their local copies.

🤖 Generated with Claude Code

Adds `slo/`, extracted from two production consumers (pay-core and blockchain-api) that
had each grown a copy of the same file. This is the shared home so they stop diverging.

  slo/burn_rate.libsonnet  the mechanics: tier table, budget sizing, PromQL conditions
  slo/rule.libsonnet       tiers + SLIs -> Grafana Unified Alerting rule groups
  slo/panels.libsonnet     the burn-rate chart and the error-budget bar gauge
  slo/tests/               promtool suites over the shipping expressions
  instant_query.libsonnet  make an alert rule's query instant, not range

The two consumers differed in exactly two things once compared, so those are the only
parameters: the label carrying the priority (`priority` vs `og_priority`) and two
sentences of annotation prose. Everything else was identical. Every knob is a hidden
field, so a consumer retunes with `burnRate + { tiers:: [...] }` rather than forking.

Also adds the repo's first CI. A shared library that nobody tests is worse than a copy
per repo, because a regression now lands everywhere at once. The suites exercise the
MECHANICS against fixture SLIs — window arithmetic, budget sizing, the noise floor,
expression splicing, the sparse-denominator hazard, and that the panels agree with the
rules about the size of the budget. Whether a given repo's objective is the right number
stays that repo's business.

Two bugs the extraction surfaced, both now pinned by fixtures:

  - A fixture whose good/bad counters share a label set made `events()` fail
    immediately. Both consumers selected the two counters with one `__name__` matcher,
    which works only because one of their metrics happens to carry an extra label:
    `increase()` drops `__name__`, so otherwise the two collapse to the same label set
    and Prometheus rejects the vector, taking every SLO rule into `exec_err_state` at
    once. The fixtures deliberately use identical label sets so nobody can reintroduce
    it.
  - `sum(good) + sum(bad)` — the other obvious form — yields nothing when either side
    has no series, which is precisely a dimension in total outage.

Existing consumers pin a submodule SHA, so nothing moves until they bump it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@chris13524
chris13524 merged commit 3754d3b into main Sep 16, 2026
1 check passed
@chris13524
chris13524 deleted the feat/slo-burn-rate branch September 16, 2026 15:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants