Server-side API

Rate limits

The limits a server-side integration runs into — per key and per endpoint — and how to back off.

Server-side requests are counted against three fixed one-minute windows.

LimitCounted perDefault
Per keyYour API key1200 / min
Per workspaceAll keys in the workspace together6000 / min
Per source IPThe address your requests come from2400 / min

These are higher than the widget's, deliberately. A browser session is one visitor; your API key carries your whole integration, including batch work.

The per-key limit is the one you will normally meet, and the one reported back to you in headers. The workspace limit only bites when several integrations share a workspace. The IP limit is counted before your key is looked up, so a flood of guessed keys from one address is stopped early — a working integration will not notice it.

Two different things count your IP

The per-source-IP limit here is a rate limit — too many requests, 429, back off. It is unrelated to the IP allowlist, which decides whether your address may use the key at all and answers 403. Same word, different mechanisms, different fixes.

Per endpoint

Inside your key's budget, every endpoint has its own limit too, so one busy route cannot use up the whole key — a runaway sync loop on one endpoint leaves the rest of your integration working.

Kind of endpointLimit, per keyWhich routes
Read600 / minEvery GET
Write300 / minEvery other method, except those below
Sensitive60 / minBooking, payment and one-time codes

The sensitive routes are the ones where a loop does more than use capacity — it creates bookings, touches payments, or sends messages to patients:

  • POST /v1/appointments/book-appointment, /book-appointment-unified, /reserve-slot, /confirm-payments
  • POST /v1/qms-appointments/pre/book, /immediate/book
  • POST /v1/session/packs, /v1/session/packs/confirm
  • POST /v1/otp/send, /v1/otp/verify
  • POST /v1/patients/send-phone-verification-otp, /verify-phone-verification-otp

A limit belongs to the endpoint, not the URL: /v1/appointments/12/cancel and /v1/appointments/13/cancel draw from the same bucket.

The one-time-code routes have their own, stricter limits on top of these — per phone number and per patient — which protect patients from being flooded with messages.

Reading the headers

Successful responses carry your remaining quota for the key, and for the endpoint you called:

X-RateLimit-Limit: 1200
X-RateLimit-Remaining: 1183
X-RateLimit-Route-Limit: 60
X-RateLimit-Route-Remaining: 57

Whichever runs out first answers the 429, and Retry-After is right either way.

When you go over

You get 429 Too Many Requests with a Retry-After in seconds:

HTTP/1.1 429 Too Many Requests
Retry-After: 37
{ "message": "Too many requests. Please slow down." }

Wait Retry-After seconds and retry. Retrying immediately does not reset the window; it only burns quota you have already spent.

Batch work belongs off the clock

If you are syncing a whole day of appointments, spread it out rather than firing 1200 requests in the first second of a minute. A short sleep between calls is far cheaper than handling 429s.

What happens if our limiter is down

The limiter fails open. If the counter store is unavailable, requests are allowed through rather than rejected — a limiter outage should not take your integration down with it.

We are stating that as a property, not an invitation. Nothing about your authentication is relaxed while it happens: your key is still checked, your address is still checked against its allowlist, and you still only reach the documented endpoints. Build to the published limits and you will never see the difference.

On this page