Skip to content

Measurement methodology

How availability, coverage, latency and incidents are computed, and what the numbers do not mean.

Updated

States

PENDING until the first evaluable observation; UP, DEGRADED and DOWN from check outcomes; UNKNOWN when observations are missing or the probe could not evaluate (after three unknown observations in a row for display); PAUSED and MAINTENANCE when you say so.

Observations

Every check records the scheduled time, actual start and finish, region, probe, configuration version, outcome, latency, status code and a bounded error. Missed execution windows are recorded as gaps and unknown observations, never as successes.

Availability

Availability is time-weighted over evaluable time, derived from recorded state transitions: (UP + DEGRADED time) / (UP + DEGRADED + DOWN time). Transition times are the observation times, so there is sampling uncertainty of up to one interval around each change. Unknown, pending and paused time is excluded and reported as coverage = evaluable / window. Maintenance time is excluded from the adjusted figure and the unadjusted figure (counting maintenance as evaluable, with downtime inside maintenance still counting as down) is also shown.

Why we do not average percentages

Two monitors with different evaluable durations cannot be averaged into one meaningful percentage, so reports list monitors individually. Likewise regional p95 values are never averaged into a global p95.

Latency

Latency is server-side request time per check. Percentiles on dashboards are nearest-rank over the raw samples of one monitor in the window. Long-range reports use merged fixed-bucket histograms and are labelled approximate.

Incidents

An incident opens after the configured number of consecutive failing observations (primary or confirmation runs) and resolves after the configured number of consecutive successes. Where a second deployed region exists and confirmation is enabled, the first failure triggers one confirmation run there; incidents confirmed from one location are labelled single probe. Flapping monitors temporarily need one extra failure and one extra success.

Notification latency

Detection latency depends on interval and timeout, confirmation adds one more check, and delivery depends on the channel. A 30-second interval is not a 30-second notification.