# Rate limits and quotas

API usage is limited in two ways: your **plan** decides whether the API is available at all, and **quota windows**
limit how many calls each scope may make per minute and per month. Separate per-address limits protect the anonymous
endpoints and the MCP server. None of this is the same as your plan's monitor and contact caps, which are described
in [Quotas and limits](/account/quotas/).

## Which plans include the API

API access is a plan feature, shown as the **API access** row in the plan comparison at
[/en/payment/packages](https://www.host-tracker.com/en/payment/packages). It is included in the **Business** and
**Enterprise** plans; check the comparison for other plans, as the lineup changes.

- Without it, the app refuses to create a token ("API access is not included in your current package, or the account
  is not active").
- `GET /account` shows `flags.apiEnabled` (and `flags.active`); `GET /account/quota` repeats `apiEnabled`. Both are
  informational - they never block a read.
- The SDKs, [ht-cli](/integrations/cli/), the [Terraform provider](/integrations/terraform/) and the
  [MCP server](/integrations/mcp/) all use the API, so they need it too.

## Quota windows

Each metered call is counted against its **scope family** and its **leaf** (a monitor list counts against `monitor`
and `monitor:read`). There is no single account-wide counter, so a monitor poller and a contact sync each get their
own allowance. Each scope has two kinds of window:

- **Per minute** - counts every request, including refused ones. It is the burst guard.
- **Per month** (30 days) - counts successful requests only; a refused call costs nothing.

The standard windows per scope:

| Plan tier | Reads per minute | Reads per month | Writes per minute | Writes per month |
|---|---|---|---|---|
| Trial | 10 | 10,000 | 5 | 500 |
| Business | 60 | 100,000 | 30 | 20,000 |
| Enterprise | 120 | 1,000,000 | 60 | 100,000 |

Instant checks (`check` scope) have their own pool with their own limits. Your account's real numbers are always what
`GET /account/quota` returns - read them rather than hard-coding this table.

## Read your headroom

`GET /account/quota` (scope `account:read`) reports what is left without spending anything:

```json
{
  "limit": 60, "used": 12, "remaining": 48, "resetAt": 1790000060,
  "pools": {
    "account": { "quotas": [ { "scope": "monitor:read", "quota": "read_per_minute", "limit": 60, "used": 12,
                               "remaining": 48, "resetAt": 1790000060, "windowSec": 60, "successOnly": false } ] },
    "check": { "quotas": [] }
  },
  "scopes": [ { "scope": "monitor:read", "description": "..." } ],
  "apiEnabled": true
}
```

The top-level `limit`/`used`/`remaining`/`resetAt` describe the tightest window. `tokenCap` appears when the calling
token has its own self-cap. An empty `quotas` list means no window is configured for that pool: calls are not
metered and carry no rate-limit headers.

## RateLimit headers

A metered call answers with:

| Header | Meaning |
|---|---|
| `RateLimit-Limit` | The size of the window that applies. |
| `RateLimit-Remaining` | Calls left in that window. |
| `RateLimit-Reset` | Seconds until the window resets. |
| `RateLimit-Policy` | Which window it is, for example `read_per_minute;q=60;w=60`. |

A `429` adds `Retry-After` (seconds).

## The two kinds of 429

| Code | Means | What to do |
|---|---|---|
| `429 quota_exceeded` | The scope's allowance for a window is spent. `errors[0]` carries `limit`, `remaining` and `resetAt` (Unix seconds). | Wait until `resetAt` (minutes for a per-minute window, days for the monthly one), spread the load, or upgrade. Do not retry in a loop. |
| `429 rate_limited` | A short-window throttle on one endpoint or one client address. `errors[0]` carries `limit`, `window` and `retryAfter`. | Wait `Retry-After` seconds and retry. |

Branch on the `code`, not the status: the SDKs and ht-cli retry `rate_limited` (and `503` with `Retry-After`)
automatically and never retry `quota_exceeded`.

## Per-address limits

- **Anonymous catalogue endpoints** (`/monitor/type`, `/contact/type`, `/agent`, `/agent/pool` and the other
  reference lists) allow 120 requests per 5 minutes per client address. `GET /agent/ip` has its own bucket of 60
  per 5 minutes. Calling them with a valid token moves you to a per-account bucket instead, which helps when many
  clients share one outgoing address.
- **MCP server:** 60 requests per minute per client address, on top of the token's quota.

## Retry safely

- Honour `Retry-After` when present; otherwise back off exponentially with jitter.
- Retry only the "wait and retry" family (`429`, `500`, `502`, `503`). A `4xx` validation or permission error fails
  the same way forever.
- Send an `Idempotency-Key` on every write so a retried write cannot run twice. Keys are remembered for 24 hours;
  the rules are in [REST API v2](/integrations/rest-api/#idempotency-key).
- For Terraform on a trial token, lower parallelism: `terraform apply -parallelism=3`.

## Related

- [REST API v2](/integrations/rest-api/)
- [API authentication, tokens and scopes](/integrations/api-authentication/)
- [Quotas and limits](/account/quotas/)
- [Error codes reference](/reference/error-codes/#rest-api-v2-error-codes)
