Burn-Rate Alerting on SLOs: Multiwindow Alerts in Prometheus That Don’t Flap

Most teams don’t have an alerting problem; they have an alerting volume problem. The on-call rotation drowns in alerts that say “CPU was high for five minutes” without saying whether anyone is actually suffering. The fix isn’t fewer alerts — it’s alerts with a mathematical contract: define what “good” means (an SLO), measure how fast you’re consuming the allowance for badness (the error budget), and page only when that consumption rate implies the budget will be gone before anyone planned for it. That’s the burn-rate alerting model, and this post shows how to implement it correctly in Prometheus and Alertmanager, including the two-window trick that kills alert flapping.

The core reference is the alerting-on-SLOs chapter of the Google SRE Workbook, and its formulas are deceptively simple. The implementation details — which windows to pair, what for, and how to keep the numbers honest — are where teams get stuck.

Error budgets and burn rates in one paragraph

An SLO says: over a compliance window (usually 30 days), at most ε of requests may fail. With a 99.9% availability SLO, ε = 0.1%, and that 0.1% is your error budget. The burn rate is how fast you’re spending it relative to the pace that exactly exhausts it at the end of the window. Burn rate 1 means you’ll have precisely zero budget left when the 30 days are up. Burn rate 14.4 means you’re spending the budget 14.4× faster than that — so fast that at a constant rate you’d exhaust a full 30-day budget in about 50 hours (30 days ÷ 14.4 ≈ 50 hours). This is the number worth alerting on, because “current error rate exceeds SLO threshold” fires constantly for systems that hover near the target, and almost never for systems that sprint to failure from a healthy baseline.

The budget consumed by the time an alert fires is simply burn rate × alerting window ÷ SLO period. Five percent of a 30-day budget spent in one hour requires a burn rate of 36. That arithmetic lets you work backwards from “how much damage am I willing to absorb before a page goes out?” to concrete thresholds.

Why single-window alerts flap (or sleep)

Suppose you alert when the 1-hour error rate exceeds 2% (a 14.4× burn rate for a 99.9% SLO). Two failure modes appear:

  • The sprint-to-failure blind spot: an incident starts as a 15% error rate for 10 minutes, then fully recovers. The 1-hour average after recovery is (0.15 × 10/60) ≈ 2.5% — the alert fires 50 minutes after the incident ended, waking someone for a problem that no longer exists.
  • The slow leak: a service runs at 2.1% errors for hours. Every hour the alert fires, gets acknowledged, self-resolves at the window boundary, and refires. The on-call learns to ignore it, and the error budget quietly empties.

Long windows miss fast incidents; short windows fire on noise. Any single window forces you to choose which failure you’ll tolerate.

Multiwindow, multi-burn-rate: the two-window pairing

The solution from the SRE Workbook is to require both a long window and a short window to exceed the same burn-rate threshold. The long window provides sensitivity (it would eventually catch the incident), and the short window provides confirmation (the error rate is still elevated right now). The canonical page-level pair for a 99.9% SLO:

  • Fast burn: 14.4× burn rate over the last 1 hour AND the last 5 minutes → page. Catches errors that would burn 2% of the 30-day budget in 1 hour.
  • Slow burn: 6× burn rate over the last 6 hours AND the last 30 minutes → page. Catches errors burning 5% of the budget in 6 hours.
  • Ticket-level (optional): 3× over 1 day AND 2× over 2 hours → file a ticket, don’t page. Significant, but not someone’s 3 a.m.

Rerun the sprint-to-failure scenario against the fast-burn rule: 15% errors for 10 minutes pushes the 5-minute window over threshold immediately, but the 1-hour window needs the elevation to persist roughly 5 minutes before its average crosses too — so the alert starts firing within minutes during the incident, and self-resolves shortly after recovery because the short window drains quickly. The slow leak trips the slow-burn rule within 30 minutes of crossing 6×, and stays firing until it’s actually fixed. Both failure modes handled, no new infrastructure.

The Prometheus implementation

Express each condition as a ratio of bad requests to total requests over a window, compare it to the burn-rate threshold, and AND the long and short windows. For an HTTP service with a 99.9% availability SLO, where “bad” means any non-2xx/3xx response, the fast-burn rule in PromQL:

groups:
  - name: slo-burn-rate
    rules:
      # Fast burn: 14.4x, windows 1h and 5m
      - alert: SLOBurnRateFast
        expr: |
          (
            sum(rate(http_requests_total{job="api", code=~"5.."}[1h]))
            /
            sum(rate(http_requests_total{job="api"}[1h]))
          ) > (14.4 * 0.001)
          and
          (
            sum(rate(http_requests_total{job="api", code=~"5.."}[5m]))
            /
            sum(rate(http_requests_total{job="api", code=~"5.."}[1h]))
          )

Wait — that second clause is wrong, and it’s the most common copy-paste bug in burn-rate alerts. The short window must compare the short-window error rate against the same threshold, not mix windows within one ratio. The correct form:

groups:
  - name: slo-burn-rate
    rules:
      - alert: SLOBurnRateFast
        expr: |
          (
            sum(rate(http_requests_total{job="api", code=~"5.."}[1h]))
              / sum(rate(http_requests_total{job="api"}[1h]))
          ) > (14.4 * 0.001)
          and
          (
            sum(rate(http_requests_total{job="api", code=~"5.."}[5m]))
              / sum(rate(http_requests_total{job="api"}[5m]))
          ) > (14.4 * 0.001)
        for: 2m
        labels:
          severity: page
        annotations:
          summary: "API burning error budget at 14.4x (99.9% SLO)"
          runbook: https://runbook.example.com/slo/fast-burn

      - alert: SLOBurnRateSlow
        expr: |
          (
            sum(rate(http_requests_total{job="api", code=~"5.."}[6h]))
              / sum(rate(http_requests_total{job="api"}[6h]))
          ) > (6 * 0.001)
          and
          (
            sum(rate(http_requests_total{job="api", code=~"5.."}[30m]))
              / sum(rate(http_requests_total{job="api"}[30m]))
          ) > (6 * 0.001)
        for: 15m
        labels:
          severity: page
        annotations:
          summary: "API burning error budget at 6x (99.9% SLO)"
          runbook: https://runbook.example.com/slo/slow-burn

Note the for: 2m and for: 15m clauses. Prometheus evaluates rules every evaluation interval, and a single noisy scrape could momentarily push a 5-minute window over threshold. The for duration demands the condition hold across consecutive evaluations before the alert actually fires — a second, cheaper denoising layer on top of the two-window logic. Keep it short (1–5 minutes for fast burn); a long for on the fast rule defeats its purpose.

If your availability SLO is user-centric rather than HTTP-status-centric — say, “99.9% of requests complete in under 300 ms” — only the numerator changes:

      - record: slo:latency_error_ratio_1h
        expr: |
          1 - (
            sum(rate(http_request_duration_seconds_count{job="api", le="0.3"}[1h]))
            /
            sum(rate(http_request_duration_seconds_count{job="api"}[1h]))
          )

The structure of the alert rules stays identical — one recording rule per window, alerts compare two of them to the same threshold. Recording rules are worth it once you have more than a couple of services: they keep the alert expressions readable and let dashboards reuse the same series.

Deriving your own numbers

The 14.4/6/3/2 thresholds aren’t sacred constants; they come from one formula. For a desired budget consumption of B% over a window of W hours with a 30-day (720-hour) SLO period:

burn rate = (B/100) × 720 ÷ W

Check: 2% over 1 hour → 0.02 × 720 = 14.4. Five percent over 6 hours → 0.05 × 720 ÷ 6 = 6. If your SLO period is a week or your tolerance for fast incidents is higher, derive your own table instead of copying the canonical one — the Workbook chapter shows the full derivation.

Keep two constraints in mind. First, the short window must be short enough that it drains quickly after recovery — a 30-minute short window means up to 30 minutes of stale confirmation after the incident ends. Second, the long windows of the fast and slow rules should differ enough (1h vs 6h) that the two alerts represent genuinely different response regimes, not the same incident paging twice.

Common pitfalls

  • Alerting on the SLO threshold itself. “Error rate > 0.1%” fires constantly for borderline services and never for sprint-to-failure ones. Always alert on burn rate, not on the raw threshold.
  • Client-side error blindness. 5xx-only SLIs miss timeouts, connection resets, and TLS failures that never produced a response. Count those in the bad numerator — usually via a request-completion counter on the client or load-balancer side.
  • Denominator collapse at low traffic. At 3 a.m. with 2 requests/minute, one failure is a 50% error rate and the fast-burn alert fires on noise. Either gate the rule with a minimum request-rate condition or accept the noise and tune for durations.
  • SLO period mismatch. The thresholds assume a 30-day window. If your compliance window is 7 days, recompute everything (720 → 168) or your burn rates are silently wrong by a factor of 4.3.
  • Budget-blind paging escalation. The burn-rate alerts tell you when to wake someone; the remaining budget tells you when to stop shipping features and fix reliability. Wire the budget numbers into your sprint planning, or the alerts become a treadmill.

Wrapping up

Burn-rate alerting converts “the system looks sick” into “we are consuming our reliability allowance at N times the sustainable pace, confirmed across two time horizons.” The Prometheus implementation is a handful of YAML — the hard part is organizational: agreeing on an SLO you actually defend, an SLI that measures user experience rather than process health, and the discipline to let the error budget arbitrate feature-versus-reliability debates. Start with one service, the 14.4×/6× pair, and a 99.9% target you believe in; the rest of the machinery generalizes from there.

Leave a Reply

Your email address will not be published. Required fields are marked *