What’s metered, and at what grain
Usage is accounted per organization. The budget is shared across every API key, every user, and the MCP connector, so spreading calls over more keys or more endpoints does not raise your ceiling — it all draws on the same organization budget. Usage is tracked along a few dimensions, each with its ownunit:
When the budget is exhausted
An exhausted budget is rejected before any expensive processing, so you are not billed for work that never ran. The response is429 Too Many Requests with:
- a
Retry-Afterheader, and - a JSON body whose
detailreports exactly what was hit and when it clears.
reset_at is when the budget refills; retry_after is the same wait in seconds and mirrors the Retry-After header. Because the budget is shared, switching to another API key does not restore access — wait for the reset.
How to behave well
- Honour
Retry-After. Back off for the stated interval rather than retrying immediately; a tight retry loop just burns the next window too. - Don’t fan out across keys to dodge the limit — it’s one organization budget.
- Prefer bulk and async endpoints over many small calls (see Imports vs bulk); they move the same data for far fewer requests.
- Poll on a sane cadence. For async jobs, every 5–10s for active jobs and 30–60s for queued ones is plenty (see Conventions ▸ Async jobs).
If enforcement is unavailable
If the usage-enforcement service itself is temporarily unavailable, the API returns503 with { "detail": { "code": "usage_enforcement_unavailable" } }. This is transient — retry after a short delay.

