# How down detection works

HostTracker is deliberately careful before it calls a monitor **Down**. One failed check is never enough: the
failure is first **re-checked from several other locations**, and only an agreed result changes the state, opens
an incident and alerts your contacts. This page follows one outage from the first failed check to recovery.

## What counts as a failed check

A single check fails when the target cannot be reached or does not pass validation: DNS cannot resolve the name
(the name does not exist), the connection is refused or times out, the TLS handshake fails a
[policy](/monitors/advanced/tls-policy/) you switched on, the HTTP status is an error, a keyword or
[assertion](/monitors/advanced/assertions/) fails, a database query or threshold fails - depending on the
[type](/monitors/monitor-types/).

A check that could not be carried out at all - the checkpoint was busy, or failed internally - is **not** a
failure of your site. It is discarded and retried, and it never changes the state.

## From failed check to Down, step by step

1. **A scheduled check fails.** It ran from one checkpoint of the monitor's
   [locations](/monitors/default-locations/). Nothing changes yet.
2. **HostTracker waits 10 seconds**, so a target that is restarting gets a moment.
3. **Up to 7 other checkpoints re-check** the target (never fewer than 3, even if some of them have to be
   less-trusted ones), launched about 2 seconds apart. They are taken from the monitor's own locations; checkpoints that recently returned bad results
   are skipped while better ones are available.
4. **The votes are counted** with the monitor's [recheck strategy](/monitors/advanced/recheck-strategy/). With
   the default **majority vote**, more Down votes than Up votes confirms the failure; a tie keeps the previous state.
5. **Confirmed: the monitor turns Down.** An [incident](/incidents/what-is-an-incident/) opens. Its start time is
   the moment of the first failed check, not the moment of confirmation, so the recorded downtime includes the
   confirmation time.
6. **Down alerts go out** to every contact subscribed to Down alerts on this monitor - immediately for contacts
   without a delay, and later for contacts with an [alert delay](/alerts/escalation/) (a contact whose delay is
   longer than the outage never hears about it).

If the re-check does **not** confirm the failure, the first result is treated as a local problem of that one
checkpoint: the state stays Up, no incident opens and no alert is sent.

### How long confirmation takes

Roughly: the failing check (up to its [timeout](/monitors/advanced/response-limits/)) + 10 seconds + the
re-check (up to the timeout again, plus about 12 seconds of stagger). A site that refuses connections is
confirmed Down within about half a minute; a site that hangs until the 40-second default timeout can take close
to two minutes.

## While the monitor stays Down

Every following check that still fails produces a **still-down** reminder (the **Repeat** event in alert
subscriptions, `repeatedlyDown` on the API) for contacts subscribed to it; how often a contact actually receives
them also depends on that contact's own settings (see [Escalation and repeats](/alerts/escalation/)). For some
types reminders are spaced out: at most once a day for DNSBL, Domain expiry and SSL certificate expiry monitors,
and once a week for a site crawl. The incident stays open and no new incident is created.

## Recovery

When a check succeeds again, the same confirmation runs in the other direction: other checkpoints re-check, and
an agreed Up closes the incident and sends the **Up** alert. The incident's end time is the moment of the first
successful check. The Up alert goes to the contacts subscribed to Up alerts who were due a Down alert for this
outage.

## Types that decide on one result

Some checks are answered by one authoritative source, so a second opinion from another checkpoint adds nothing.
For these, the first result is final and there is no recheck:

- **Database**, **SNMP** and **Counter** (they run on HostTracker's internal network)
- **Domain expiry** and **Web Risk**
- **DNSBL** (a new listing is confirmed by a second lookup inside the check itself)
- **Site crawl** (one run is hundreds of page fetches; its verdict is already an aggregate)

**SSL/TLS certificate expiry** does use the recheck. **Web content check** monitors always use the majority vote;
their strategy cannot be changed.

## Special cases

- **DNS resolver trouble.** If checkpoints cannot resolve the host because *their* DNS resolver fails (a server
  failure or timeout rather than "this name does not exist"), the result is discarded and retried elsewhere.
  When three different checkpoints in a row report the same resolver failure, it is recorded as a real Down
  ("DNS resolution failed (confirmed by 3 agents)"). A name that does not exist (NXDOMAIN) is a normal Down result
  and goes through the usual recheck.
- **Alerts held on low-confidence evidence.** If every result that agreed on a change came from checkpoints that
  had to fall back to a secondary DNS resolver, the state still changes but the alert for that change is held
  back.
- **Paused monitors** are not checked at all, so they can neither go Down nor recover. See
  [Pause, enable and delete](/monitors/pause-enable-delete/).

## Maintenance windows

During a [maintenance window](/maintenance/overview/) the monitor keeps being checked, but:

- it shows as `maintenance` rather than `down` on the dashboard and the API;
- with alert suppression on, Down, Up and still-down alerts are not sent - and if the monitor is **still Down
  when the window ends**, one Down alert is sent at that moment;
- with statistics suppression on, the time inside the window is left out of the uptime percentage.

See [What maintenance suppresses](/maintenance/what-it-suppresses/).

## See it and verify it

| What you want | App | API (scope `monitor:read`) | MCP |
|---|---|---|---|
| Current state and since when | The monitor's row on **Sites** | `GET /monitor/{id}` -> `state`, `since` | `get_monitor` |
| The last up/down change | The row's incident panel | `GET /monitor/{id}?expand=lastIncident` | `get_monitor` with `expand=lastIncident` |
| All incidents | The incidents panel on **Sites** | `GET /monitor/{id}/incident`, `GET /monitor/incident/{id}` | `list_incidents`, `get_incident` |
| Which checkpoints voted, and how | The incident's checks | `GET /monitor/incident/{id}/check`, `GET /monitor/{id}/result/{resultId}?expand=recheck` | `list_monitor_results` with `expand=recheck` |

## Related

- [Recheck strategy](/monitors/advanced/recheck-strategy/)
- [Choose monitoring locations](/monitors/default-locations/)
- [Why is my monitor down?](/incidents/why-is-it-down/)
- [Short outages recorded as long downtime](/troubleshooting/short-outages/)
