Per-API-key usage limits and spend alerts
Every API key on your workspace now reports its own live usage, and you can bound each key independently: a requests-per-minute rate limit, a monthly request quota, and a usage-alert threshold that flags a key before it exhausts its quota. One noisy integration — or a leaked key — trips its own limit instead of spending the throughput your other keys depend on. This is the per-key scoping mature communications APIs expose: the org-wide default applies to every key until you override a single key, and you can read live usage and headroom for each key at any time.Rate limits and monthly quotas
There are two independent controls per key — set either, both, or none:- Rate limit (requests per minute). When a key crosses its per-minute window, that key alone receives
429 RATE_LIMIT_EXCEEDEDwith aRetry-Afterhint. Your other keys are untouched. - Monthly request quota. A hard ceiling per UTC month. Cross it and the key gets
429with a monthly-quota code until the month resets on the 1st (UTC).
The usage-alert threshold
Set a whole-percent alert threshold (1–100) to flag any key whose monthly-quota usage crosses it.GET /api/v1/developer/governance marks each key with alerting: true once its usage reaches the threshold, so a runaway key shows up in your console (and your polling) before it hits the hard quota.
The threshold is advisory — it does not block traffic. Blocking happens only at the rate limit or the quota.
Read state and live usage
Owner / admin / developer scope.Polyglot samples
The sameGET /api/v1/developer/governance in the Node, Python, and Go SDKs. Governance is not on the SDK resource list, so each SDK exposes it through its low-level request escape hatch — which still handles auth, JSON serialisation, and 429/5xx retry the same way every resource does.
Node SDK
Python SDK
Go SDK
org_default vs per-key override
The response carries two scopes side by side.org_default is the baseline applied to every key that lacks an override; a key’s own override replaces those dimensions when present. effective_limits is the already-resolved value the guard enforces — per-key override for a dimension when one exists, otherwise the org_default for that dimension — so read it instead of re-deriving the precedence.
alert_threshold_percent is org-wide only — it is not a per-key field — so it lives at the top level, not inside org_default. bounds is the platform ceiling each value must stay under; a value above it is rejected at config time, not silently clamped.
Response (one key shown):
effective_limitsis what actually gates the key — the per-key override when one exists, otherwise the org default.overrideisnullwhen the key inherits the org default.monthly_quotaisnullwhen neither the key nor the org has a monthly quota set (usage is still reported; nothing blocks).usage_percentis floored — 79.9% reports79and does not prematurely trip an 80% threshold.- The read hydrates at most 250 keys with live usage;
keys_truncated: truetells you the list was cut.
Set the org-wide default
null on a field to clear it (the key inherits the platform default). alert_threshold_percent accepts an integer from 1 to 100, or null to disable flagging.
The body must contain at least one field — an empty body is a 422, not a silent no-op.
Override a single key
null on a dimension to inherit the org default for that dimension; send both as null to remove the override entirely. A key you do not own returns 404 — the write can never be applied to another tenant’s key.
429 envelope on the retry path
A key that crosses its per-minute window or its monthly quota gets a429 with a stable error code you branch on — never the human-readable message. The two codes are distinct so you can tell a transient rate window from a month-long quota ceiling:
RATE_LIMIT_EXCEEDED— the key crossed its per-minute window. Wait out theRetry-Afterhint (60 seconds) and retry the same request; the window resets on its own.QUOTA_EXCEEDED— the key exhausted its monthly request quota. The quota resets at the start of next month (UTC); raise the quota (per-key override or the org default) or the next request is rejected too.
Rate-limit 429 (retryable)
Retry-After header and error.details.retry_after carry the same value (60 seconds) — read either and retry once it elapses. The schema is the standard error envelope: error.code is the decision key, error.status is the HTTP class, and error.details carries the structured context (limit, used, window, retry_after).
Monthly-quota 429 (not retryable this month)
The same 429 status, a different code:QUOTA_EXCEEDED has no retry_after — the quota does not reset until the 1st of next month (UTC), so retrying immediately cannot help. Raise the quota via PUT /api/v1/developer/governance/org or PUT /api/v1/developer/governance/keys/:keyId and the very next request succeeds; there is no cache TTL to wait out.
Branch on the code in the SDK
The Node SDK surfaces both asOrbitRateLimitError; branch on .code to split the retryable rate window from the monthly ceiling:
Node SDK
Python SDK
Go SDK
Constraints and behavior
- Writes take effect on the next request — there is no propagation lag.
- Both writes are audit-logged (
developer.api_governance.org_updated,developer.api_governance.key_updated). - Existing keys are unchanged until you set a limit — a key left unconfigured behaves exactly as before.
- Values of
0are rejected (a zero cap would hard-block the key, which is never the intent; to disable a key, revoke it). Usenullto clear instead. - The rate/quota guards are fail-open on infrastructure errors: a cache blip never 429s or blocks a well-behaved key; only a real over-limit verdict does.