Skip to content

Rate limits and quotas

View as Markdown

API usage is limited in two ways: your plan decides whether the API is available at all, and quota windows limit how many calls each scope may make per minute and per month. Separate per-address limits protect the anonymous endpoints and the MCP server. None of this is the same as your plan’s monitor and contact caps, which are described in Quotas and limits.

API access is a plan feature, shown as the API access row in the plan comparison at /en/payment/packages. It is included in the Business and Enterprise plans; check the comparison for other plans, as the lineup changes.

  • Without it, the app refuses to create a token (“API access is not included in your current package, or the account is not active”).
  • GET /account shows flags.apiEnabled (and flags.active); GET /account/quota repeats apiEnabled. Both are informational - they never block a read.
  • The SDKs, ht-cli, the Terraform provider and the MCP server all use the API, so they need it too.

Each metered call is counted against its scope family and its leaf (a monitor list counts against monitor and monitor:read). There is no single account-wide counter, so a monitor poller and a contact sync each get their own allowance. Each scope has two kinds of window:

  • Per minute - counts every request, including refused ones. It is the burst guard.
  • Per month (30 days) - counts successful requests only; a refused call costs nothing.

The standard windows per scope:

Plan tier Reads per minute Reads per month Writes per minute Writes per month
Trial 10 10,000 5 500
Business 60 100,000 30 20,000
Enterprise 120 1,000,000 60 100,000

Instant checks (check scope) have their own pool with their own limits. Your account’s real numbers are always what GET /account/quota returns - read them rather than hard-coding this table.

GET /account/quota (scope account:read) reports what is left without spending anything:

{
"limit": 60, "used": 12, "remaining": 48, "resetAt": 1790000060,
"pools": {
"account": { "quotas": [ { "scope": "monitor:read", "quota": "read_per_minute", "limit": 60, "used": 12,
"remaining": 48, "resetAt": 1790000060, "windowSec": 60, "successOnly": false } ] },
"check": { "quotas": [] }
},
"scopes": [ { "scope": "monitor:read", "description": "..." } ],
"apiEnabled": true
}

The top-level limit/used/remaining/resetAt describe the tightest window. tokenCap appears when the calling token has its own self-cap. An empty quotas list means no window is configured for that pool: calls are not metered and carry no rate-limit headers.

A metered call answers with:

Header Meaning
RateLimit-Limit The size of the window that applies.
RateLimit-Remaining Calls left in that window.
RateLimit-Reset Seconds until the window resets.
RateLimit-Policy Which window it is, for example read_per_minute;q=60;w=60.

A 429 adds Retry-After (seconds).

Code Means What to do
429 quota_exceeded The scope’s allowance for a window is spent. errors[0] carries limit, remaining and resetAt (Unix seconds). Wait until resetAt (minutes for a per-minute window, days for the monthly one), spread the load, or upgrade. Do not retry in a loop.
429 rate_limited A short-window throttle on one endpoint or one client address. errors[0] carries limit, window and retryAfter. Wait Retry-After seconds and retry.

Branch on the code, not the status: the SDKs and ht-cli retry rate_limited (and 503 with Retry-After) automatically and never retry quota_exceeded.

  • Anonymous catalogue endpoints (/monitor/type, /contact/type, /agent, /agent/pool and the other reference lists) allow 120 requests per 5 minutes per client address. GET /agent/ip has its own bucket of 60 per 5 minutes. Calling them with a valid token moves you to a per-account bucket instead, which helps when many clients share one outgoing address.
  • MCP server: 60 requests per minute per client address, on top of the token’s quota.
  • Honour Retry-After when present; otherwise back off exponentially with jitter.
  • Retry only the “wait and retry” family (429, 500, 502, 503). A 4xx validation or permission error fails the same way forever.
  • Send an Idempotency-Key on every write so a retried write cannot run twice. Keys are remembered for 24 hours; the rules are in REST API v2.
  • For Terraform on a trial token, lower parallelism: terraform apply -parallelism=3.