How down detection works
HostTracker is deliberately careful before it calls a monitor Down. One failed check is never enough: the failure is first re-checked from several other locations, and only an agreed result changes the state, opens an incident and alerts your contacts. This page follows one outage from the first failed check to recovery.
What counts as a failed check
Section titled “What counts as a failed check”A single check fails when the target cannot be reached or does not pass validation: DNS cannot resolve the name (the name does not exist), the connection is refused or times out, the TLS handshake fails a policy you switched on, the HTTP status is an error, a keyword or assertion fails, a database query or threshold fails - depending on the type.
A check that could not be carried out at all - the checkpoint was busy, or failed internally - is not a failure of your site. It is discarded and retried, and it never changes the state.
From failed check to Down, step by step
Section titled “From failed check to Down, step by step”- A scheduled check fails. It ran from one checkpoint of the monitor’s locations. Nothing changes yet.
- HostTracker waits 10 seconds, so a target that is restarting gets a moment.
- Up to 7 other checkpoints re-check the target (never fewer than 3, even if some of them have to be less-trusted ones), launched about 2 seconds apart. They are taken from the monitor’s own locations; checkpoints that recently returned bad results are skipped while better ones are available.
- The votes are counted with the monitor’s recheck strategy. With the default majority vote, more Down votes than Up votes confirms the failure; a tie keeps the previous state.
- Confirmed: the monitor turns Down. An incident opens. Its start time is the moment of the first failed check, not the moment of confirmation, so the recorded downtime includes the confirmation time.
- Down alerts go out to every contact subscribed to Down alerts on this monitor - immediately for contacts without a delay, and later for contacts with an alert delay (a contact whose delay is longer than the outage never hears about it).
If the re-check does not confirm the failure, the first result is treated as a local problem of that one checkpoint: the state stays Up, no incident opens and no alert is sent.
How long confirmation takes
Section titled “How long confirmation takes”Roughly: the failing check (up to its timeout) + 10 seconds + the re-check (up to the timeout again, plus about 12 seconds of stagger). A site that refuses connections is confirmed Down within about half a minute; a site that hangs until the 40-second default timeout can take close to two minutes.
While the monitor stays Down
Section titled “While the monitor stays Down”Every following check that still fails produces a still-down reminder (the Repeat event in alert
subscriptions, repeatedlyDown on the API) for contacts subscribed to it; how often a contact actually receives
them also depends on that contact’s own settings (see Escalation and repeats). For some
types reminders are spaced out: at most once a day for DNSBL, Domain expiry and SSL certificate expiry monitors,
and once a week for a site crawl. The incident stays open and no new incident is created.
Recovery
Section titled “Recovery”When a check succeeds again, the same confirmation runs in the other direction: other checkpoints re-check, and an agreed Up closes the incident and sends the Up alert. The incident’s end time is the moment of the first successful check. The Up alert goes to the contacts subscribed to Up alerts who were due a Down alert for this outage.
Types that decide on one result
Section titled “Types that decide on one result”Some checks are answered by one authoritative source, so a second opinion from another checkpoint adds nothing. For these, the first result is final and there is no recheck:
- Database, SNMP and Counter (they run on HostTracker’s internal network)
- Domain expiry and Web Risk
- DNSBL (a new listing is confirmed by a second lookup inside the check itself)
- Site crawl (one run is hundreds of page fetches; its verdict is already an aggregate)
SSL/TLS certificate expiry does use the recheck. Web content check monitors always use the majority vote; their strategy cannot be changed.
Special cases
Section titled “Special cases”- DNS resolver trouble. If checkpoints cannot resolve the host because their DNS resolver fails (a server failure or timeout rather than “this name does not exist”), the result is discarded and retried elsewhere. When three different checkpoints in a row report the same resolver failure, it is recorded as a real Down (“DNS resolution failed (confirmed by 3 agents)”). A name that does not exist (NXDOMAIN) is a normal Down result and goes through the usual recheck.
- Alerts held on low-confidence evidence. If every result that agreed on a change came from checkpoints that had to fall back to a secondary DNS resolver, the state still changes but the alert for that change is held back.
- Paused monitors are not checked at all, so they can neither go Down nor recover. See Pause, enable and delete.
Maintenance windows
Section titled “Maintenance windows”During a maintenance window the monitor keeps being checked, but:
- it shows as
maintenancerather thandownon the dashboard and the API; - with alert suppression on, Down, Up and still-down alerts are not sent - and if the monitor is still Down when the window ends, one Down alert is sent at that moment;
- with statistics suppression on, the time inside the window is left out of the uptime percentage.
See What maintenance suppresses.
See it and verify it
Section titled “See it and verify it”| What you want | App | API (scope monitor:read) |
MCP |
|---|---|---|---|
| Current state and since when | The monitor’s row on Sites | GET /monitor/{id} -> state, since |
get_monitor |
| The last up/down change | The row’s incident panel | GET /monitor/{id}?expand=lastIncident |
get_monitor with expand=lastIncident |
| All incidents | The incidents panel on Sites | GET /monitor/{id}/incident, GET /monitor/incident/{id} |
list_incidents, get_incident |
| Which checkpoints voted, and how | The incident’s checks | GET /monitor/incident/{id}/check, GET /monitor/{id}/result/{resultId}?expand=recheck |
list_monitor_results with expand=recheck |

