Measurement methodology
How availability, coverage, latency and incidents are computed, and what the numbers do not mean.
Updated
States
PENDING until the first evaluable observation; UP, DEGRADED and DOWN from check outcomes; UNKNOWN when observations are missing or the probe could not evaluate (after three unknown observations in a row for display); PAUSED and MAINTENANCE when you say so.
Observations
Every check records the scheduled time, actual start and finish, region, probe, configuration version, outcome, latency, status code and a bounded error. Missed execution windows are recorded as gaps and unknown observations, never as successes.
Availability
Availability is time-weighted over evaluable time, derived from recorded state transitions: (UP + DEGRADED time) / (UP + DEGRADED + DOWN time). Transition times are the observation times, so there is sampling uncertainty of up to one interval around each change. Unknown, pending and paused time is excluded and reported as coverage = evaluable / window. Maintenance time is excluded from the adjusted figure and the unadjusted figure (counting maintenance as evaluable, with downtime inside maintenance still counting as down) is also shown.
Why we do not average percentages
Two monitors with different evaluable durations cannot be averaged into one meaningful percentage, so reports list monitors individually. Likewise regional p95 values are never averaged into a global p95.
Latency
Latency is server-side request time per check. Percentiles on dashboards are nearest-rank over the raw samples of one monitor in the window. Long-range reports use merged fixed-bucket histograms and are labelled approximate.
Incidents
An incident opens after the configured number of consecutive failing observations (primary or confirmation runs) and resolves after the configured number of consecutive successes. Where a second deployed region exists and confirmation is enabled, the first failure triggers one confirmation run there; incidents confirmed from one location are labelled single probe. Flapping monitors temporarily need one extra failure and one extra success.
Notification latency
Detection latency depends on interval and timeout, confirmation adds one more check, and delivery depends on the channel. A 30-second interval is not a 30-second notification.