Skip to content

Uptime percentage and SLA

View as Markdown

Uptime percentage answers “what share of the time was this monitor up” over a window you choose. An SLA target is the percentage you promise (for example 99.9%), and HostTracker shows whether you met it and how much downtime you still had room for.

For a monitor and a time window:

uptime % = time up / (time measured - downtime inside maintenance windows)
  • Time measured is the time the monitor was actually being checked in the window. Time before the monitor existed and time it was paused are not counted at all.
  • Time up runs from each confirmed recovery to the next confirmed outage; downtime runs from a confirmed down transition to the confirmed recovery - the length of each incident.
  • Maintenance: downtime inside a maintenance window that suppresses Stats is left out of the calculation entirely - it counts as neither up nor down. Up time inside such a window still counts as up. A window that suppresses only alerts does not change the numbers.

The API returns the parts separately, so you can see exactly how a figure was made:

Member (GET /monitor/result/summary) Meaning
totalSec Time measured in the window.
upSec, downSec Time up and time down.
maintenance.upSec, maintenance.downSec The part of those that fell inside maintenance windows.
downSpans How many down periods.
uptimePercent The result. Use it as given rather than recomputing it.
slaTarget, slaMet, errorBudgetSecRemaining The SLA view, when a target applies.

Rows come back ordered by monitor id, and the endpoint has no sort parameter; to list monitors worst first, sort the rows by uptimePercent yourself (rows without it had no data) - see Sort a summary, worst first.

Why the same monitor shows different percentages

Section titled “Why the same monitor shows different percentages”

Each figure is its own calculation over its own window: today, the last 7 days, this month and this year are four different numbers, and a single hour of downtime moves a daily figure far more than a yearly one. A 30-day window has 43,200 minutes, so each minute of downtime costs about 0.0023%.

Very short outages can look longer than they felt: an outage lasts from the check that confirmed it to the check that confirmed recovery, so the check interval and recheck time set the resolution. See Short outages recorded as long downtime.

You can set a target in two places:

  • On a monitor - the monitor’s slaTarget field (set it through the API, for example PATCH /monitor/{id} with {"slaTarget": 99.9}). Uptime summaries for that monitor are measured against it.
  • On a status page - the SLA target on the page’s Settings tab (settings.slaTarget, Webmaster band and above). The public page then shows the uptime “against a 99.9% target”, an error budget for the last 30 days (for example “1 h 12 min used of 43 min allowed”), and a 12-month compliance table visitors can expand and export as CSV, also published as sla-export.json - see Embed and export.

The error budget is the downtime a target allows over a period: (100% - target) x period. At 99.9% that is about 43 minutes in 30 days; at 99.99%, about 4 minutes.

For a one-off comparison, GET /monitor/result/summary?...&sla=99.95 measures against a target you pass for that request only.

  1. GET /monitor/result/summary?monitor=ID&from=START&to=END for the totals.
  2. GET /monitor/incident?monitor=ID&from=START&to=END for the outages behind downSec - their durationSec values add up to the downtime (an incident that started before the window counts only its part inside it).
  3. GET /maintenance?monitor=ID&from=START&to=END for the windows behind maintenance.

MCP: get_uptime_summary, list_incidents, list_maintenance.