Rate limits
The limits a server-side integration runs into — per key and per endpoint — and how to back off.
Server-side requests are counted against three fixed one-minute windows.
| Limit | Counted per | Default |
|---|---|---|
| Per key | Your API key | 1200 / min |
| Per workspace | All keys in the workspace together | 6000 / min |
| Per source IP | The address your requests come from | 2400 / min |
These are higher than the widget's, deliberately. A browser session is one visitor; your API key carries your whole integration, including batch work.
The per-key limit is the one you will normally meet, and the one reported back to you in headers. The workspace limit only bites when several integrations share a workspace. The IP limit is counted before your key is looked up, so a flood of guessed keys from one address is stopped early — a working integration will not notice it.
Two different things count your IP
The per-source-IP limit here is a rate limit — too many requests, 429,
back off. It is unrelated to the
IP allowlist, which decides whether your
address may use the key at all and answers 403. Same word, different
mechanisms, different fixes.
Per endpoint
Inside your key's budget, every endpoint has its own limit too, so one busy route cannot use up the whole key — a runaway sync loop on one endpoint leaves the rest of your integration working.
| Kind of endpoint | Limit, per key | Which routes |
|---|---|---|
| Read | 600 / min | Every GET |
| Write | 300 / min | Every other method, except those below |
| Sensitive | 60 / min | Booking, payment and one-time codes |
The sensitive routes are the ones where a loop does more than use capacity — it creates bookings, touches payments, or sends messages to patients:
POST /v1/appointments/book-appointment,/book-appointment-unified,/reserve-slot,/confirm-paymentsPOST /v1/qms-appointments/pre/book,/immediate/bookPOST /v1/session/packs,/v1/session/packs/confirmPOST /v1/otp/send,/v1/otp/verifyPOST /v1/patients/send-phone-verification-otp,/verify-phone-verification-otp
A limit belongs to the endpoint, not the URL: /v1/appointments/12/cancel and
/v1/appointments/13/cancel draw from the same bucket.
The one-time-code routes have their own, stricter limits on top of these — per phone number and per patient — which protect patients from being flooded with messages.
Reading the headers
Successful responses carry your remaining quota for the key, and for the endpoint you called:
X-RateLimit-Limit: 1200
X-RateLimit-Remaining: 1183
X-RateLimit-Route-Limit: 60
X-RateLimit-Route-Remaining: 57Whichever runs out first answers the 429, and Retry-After is right either way.
When you go over
You get 429 Too Many Requests with a Retry-After in seconds:
HTTP/1.1 429 Too Many Requests
Retry-After: 37{ "message": "Too many requests. Please slow down." }Wait Retry-After seconds and retry. Retrying immediately does not reset the
window; it only burns quota you have already spent.
Batch work belongs off the clock
If you are syncing a whole day of appointments, spread it out rather than firing 1200 requests in the first second of a minute. A short sleep between calls is far cheaper than handling 429s.
What happens if our limiter is down
The limiter fails open. If the counter store is unavailable, requests are allowed through rather than rejected — a limiter outage should not take your integration down with it.
We are stating that as a property, not an invitation. Nothing about your authentication is relaxed while it happens: your key is still checked, your address is still checked against its allowlist, and you still only reach the documented endpoints. Build to the published limits and you will never see the difference.