Skip to content

Error budgets and burn-rate alerts in practice

How UptimeMonitor360 computes an SLO's error budget from time-weighted availability, what the burn rate means, when a burn alert fires and how to read the budget chart.

By the UptimeMonitor360 team · updated

The budget

An SLO objective names a monitor, a target such as 99.9% and a window (7, 28, 30 or 90 days, or a calendar month). The error budget is the window length multiplied by one minus the target: for 99.9% over 30 days, 43.2 minutes of allowed downtime. The budget is spent by minutes in which the monitor was down (and, when the objective says so, degraded); unknown minutes, maintenance and periods without observations are excluded from both numerator and denominator and reported as reduced coverage.

Where the numbers come from

Budgets are computed from the hourly and daily rollups that also drive the availability charts, so the SLO page and the monitor page always agree. Windows up to 31 days use hourly rollups; longer windows use daily rollups with hourly precision at the edges.

Burn rate

The burn rate compares the share of the budget already consumed with the share of the window already elapsed. A rate of 1.0 means the budget will run out exactly at the end of the window; 2.0 means it runs out halfway. A short outage early in a window shows a high rate because little of the window has elapsed; the projected exhaustion day on the chart puts it in perspective.

Burn alerts

When burn alerts are enabled on an objective, the worker evaluates every objective every five minutes and sends an alert through the notification rules subscribed to the SLO burn event when the burn rate exceeds the configured threshold. A six-hour cooldown per objective keeps the alert from repeating while the rate stays high. The alert carries the current availability, the remaining budget and the rate, so the decision to pause releases can be made from the message.

Reading the chart

The budget chart shows consumed budget against the ideal even-burn line. A curve above the line means faster-than-sustainable consumption. The chart marks the incidents that consumed budget, so a single large incident is distinguishable from many small ones, which usually calls for a different remediation.

Latency objectives

An objective can also target a latency percentile, such as p95 under 800 ms. Each hour counts as good when the percentile computed from the latency histogram meets the target. This uses the same histograms that produce the latency charts, so it reflects what visitors experienced from the probe's location, not an average.

Monitor the behaviour described here continuously: start free with 5 monitors or try the free tools.