Skip to content

What is an incident

View as Markdown

An incident is a confirmed down episode of one monitor: it opens when HostTracker confirms the monitor is down and closes when the monitor is confirmed up again. Incidents are what your downtime, uptime percentages, reports and status page history are built from. HostTracker creates them automatically; you can add a note to explain what happened.

  1. A check fails. Before believing it, HostTracker rechecks from other locations according to the monitor’s recheck strategy.
  2. When the failure is confirmed, the monitor changes to down and an incident opens. Its start is the moment of that confirmed down transition, and its cause is the error the confirming checks reported.
  3. While it stays down, further failing checks are added to the incident.
  4. When a check confirms the monitor is up again, the incident closes. Its end is the recovery moment.

A single flaky check from one location that the recheck does not confirm never becomes an incident. An incident is also recorded during a maintenance window - it is marked as under maintenance and no alert is sent for it.

Incidents are not the same as the announcements you post on a status page: an incident is recorded by the checking engine; a status page incident is what you tell your visitors.

Field API member Meaning
Id id An opaque id, used by GET /monitor/incident/{id}.
Monitor monitorId (and monitor with expand=monitor) The monitor it belongs to.
Start start When the confirmed down episode began (Unix seconds).
End end When it resolved. For an open incident, the last time it was observed down.
Duration durationSec End minus start, in seconds.
State state open or resolved.
Severity severity From the duration: minor (under 5 minutes), major (5 minutes to 1 hour), critical (1 hour or more).
Cause cause The error that opened it: type, code, message, description and a stable codename such as Timeout or DnsHostNotFound - see error codes.
Under maintenance underMaintenance True when it began inside a maintenance window.
Failed checks checkCount How many failing checks were recorded in it.
Comment comment Your note (empty when there is none).
Recheck recheck (with expand=recheck) Which location detected the failure, which locations confirmed it (with their errors) and which still saw it up.
Timeline timeline (single read) The check events that opened and closed it.
  • A monitor’s statistics page (/sites/stats/{id}): the Latest incidents card lists the last 12 months, each with its state (Resolved or Ongoing), cause, start and duration. The Outages view of Recent checks shows outages in the selected period. Click one to open its detail panel.
  • The uptime report (/sites/uptime): the Incidents tab lists incidents across every monitor - see The account-wide uptime report.
  • Your status pages, as monitoring-detected outages in the history (if that display switch is on).
  • Alerts and webhooks: the down and up alerts, and the incident.opened / incident.closed webhook events.

A comment explains an incident for later - the root cause, the fix, a ticket number.

  1. Open the incident from Latest incidents or the Outages list on the monitor’s statistics page.
  2. In the detail panel, click Add note (or type in Add a note about this outage…).
  3. Save. To change it, edit the note; to remove it, delete it.

Each incident has one comment. Saving a new one replaces the previous one.

API: POST /monitor/incident/{id}/comment with {"comment": "Root cause: expired DB credentials"} (scope monitor:write). The comment member is required; "" clears it. The answer is the updated incident. MCP: comment_incident.

Task Operation MCP tool
Incidents across the account or selected monitors GET /monitor/incident list_incidents
One monitor’s incidents GET /monitor/{monitorId}/incident list_incidents with monitor
One incident with its timeline GET /monitor/incident/{id} get_incident
The failing checks inside one incident GET /monitor/incident/{id}/check api_request

Filters and examples: Reading results and incidents.