Tenant isolation
Orbit is a multi-tenant platform. Every organization’s working data lives in its own PostgreSQL schema namedtenant_<tenant_id> (dashes become
underscores in the schema name). Shared catalog and routing tables —
the things every tenant needs to read but no tenant owns — live in the
public schema. This page defines the terms, catalogues both splits, and
walks the request-resolution chain end to end so you can reconstruct it
from the code.
What a tenant is
Three related terms appear across the docs:- Organization (org) — a row in the shared
public.organizationscatalog with its ownid,tenant_id, and provisioning status (tenant_schema_ready). Your API key is issued against one org. - Tenant schema — one PostgreSQL schema (
tenant_<tenant_id>) holding that org’s working data: contacts, messages, campaigns, agents, sender pools, and so on. Roughly 200 tables form this per-tenant catalog. - Subaccount — a separate org (row) under a reseller parent that shares the parent’s tenant schema. Sibling subaccounts can co-own one schema; ownership gating decides what each can see.
tenant_schema_ready flag flips. On
most endpoints a not-yet-ready schema returns 503; endpoints that
explicitly tolerate provisioning-in-progress opt out.
Section 1 — The public schema: catalog tables every tenant reads
public holds tables every tenant reads, or tables that identify tenants
themselves. Reading a row here doesn’t read another tenant’s data — the
org boundary is still enforced by ownership columns. Mutations on a public
table are gated the same way: the platform IP denylist, the DNIS routing
table, and every per-org row in organizations all carry an org-ownership
check on write. public is catalog-only: no per-tenant working data
(messages, contacts, recordings) ever lives there.
These are catalog or routing tables: the shape is “one row per org or
one row per platform artifact”, never “one table per tenant”.
Section 2 — Tenant schemas: one per tenant, named tenant_<tenant_id>
Everything an org creates, sends, or receives lives in its own tenant
schema. Senders, inbound SMS routing, SMPP credentials, sender pools,
contacts, campaigns, flows, agent state, recordings, inbox, CDP events,
WFM — all are per-tenant tables. Provisioning creates the schema, builds
the full base table set synchronously, and flips tenant_schema_ready;
the tenant migration set then reconciles any drift in the background
(Section 2.1).
Because each tenant schema carries its own copy of the same table set, a
pool id — or any other per-tenant resource id — is only unique within
one tenant schema. sender_pools ids repeat across tenants; Postgres
enforces uniqueness per-tenant via the table’s primary key inside each
schema. A sibling subaccount shares your tenant schema, so a number it
owns is visible to the full org-family tree; a sender pool it owns is
visible too, but only inside your family’s schema — never across tenants.
Section 2.1 — How tenant schemas are provisioned and migrated
A schema is not justCREATE SCHEMA: it must carry the ~200-table set,
stay consistent with every other tenant’s copy of that set, and report
its readiness through one flag (organizations.tenant_schema_ready).
The mechanism is split in two so the customer-facing path never waits on
the slow path.
Provisioning is a two-phase split — fast core, deferred catch-up.
Signup and every other provisioning path run
ensureTenantSchema(url, tenantId) (packages/database/src/tenant.ts)
synchronously: inside one transaction it CREATE SCHEMA IF NOT EXISTS,
SETs search_path, and builds the full canonical table set — every
base table a customer-facing API touches exists the moment it returns —
and the provisioning path flips tenant_schema_ready = true. That is a
fast-core path:
account-scoped APIs stop returning 503 within seconds of signup instead
of waiting for the full migration walk. The slower catch-up —
runSingleTenantMigrations(tenantId)
(packages/database/src/tenant-migrate.ts), dispatched fire-and-forget
by apps/api/src/lib/provision-tenant-migrations.ts
— is an idempotent no-op on a fresh schema: every migration body is
guarded by IF NOT EXISTS / to_regclass, so re-running it reconciles
any column an older migration added without ever blocking the request.
Migrations are versioned per tenant and drift-repaired on three
tracks. The tenant migration set (TENANT_MIGRATIONS in
packages/database/src/tenant-migrate.ts) is one composed array — every
public-schema migration in the repo plus every tenant-schema migration —
dispatched one body at a time; only bodies that reach the tenant schema
run per-tenant (Section 1’s catalog tables are built there by the
public-schema bodies in the same set). Each tenant schema carries a
_migrations ledger recording which numbered migration ids have been
applied to that schema. New platform releases append new numbered
tenant migrations; the periodic runTenantMigrations() sweep
(packages/database/src/tenant-migrate.ts) enumerates every tenant and
applies only the ones each schema’s ledger has not recorded, so a
long-lived tenant and a same-day signup converge to the same shape
without the sweep ever re-running already-applied bodies. Migration
order is fixed by the array — arrivals converge to release order, never
to per-tenant signup order. Because the sweep is periodic, a newly
released migration always has an until-next-sweep window in which some
tenant schemas genuinely lack its table or column; the read-path and
write-path guards below exist to cover exactly that window, and every
schema change has to be tolerant of it during rollout.
- Read-path:
withTenantSchemaGuard(apps/api/src/lib/with-tenant-schema-guard.ts, Section 4’s attach path) converts a missing-table/missing-column error on a not-yet-migrated tenant into a structuredTENANT_SCHEMA_INCOMPLETE503 (or an empty degrade on read-only endpoints) rather than a wrong-tenant read — and the auth middleware’shealTenantSchemaIfNeeded(packages/auth/src/middleware.ts) re-runs the same idempotentensureTenantSchemain-band on the first request that finds an org stuck attenant_schema_ready = false. - Write-path: the 15-minute
scheduler-tenant-schema-repair(apps/webhook-worker/src/scheduler-tenant-schema-repair.ts) sweeps for the reverse drift — an org whose flag is alreadytruebut whose schema is missing — and re-provisions it. The per-request self-heal only fires when the flag isfalse, so this scheduler is the repair loop for the “flag ready, schema gone” case no request ever triggers. - Reconciliation sweep: nobody has to verify a background apply
dropped a tenant.
runTenantMigrationsInBackground’s own failure never rolls the org back: the periodicrunTenantMigrations()sweep and the per-request self-heal both re-run the same idempotent pass, so a dropped or partial apply converges on the next sweep with no manual step.
tenant_schema_ready and the 503 boundary. Until provisioning
completes, most endpoints return 503 instead of querying a half-built
schema; the flag is the single readiness signal every gate (Section 4),
scheduler fan-out, and repair loop keys off. Two enumerations keep that
cheap at scale: scheduler fan-outs read the ready flag from
organizations through the short-TTL-cached, single-flight
getAllTenantIds (apps/webhook-worker/src/lib/get-all-tenant-ids.ts)
rather than probing each tenant schema, and the runTenantMigrations()
sweep resolves every candidate schema’s existence — and which schemas
carry a _migrations table — in two batched information_schema
queries for the whole fleet, O(1) round trips instead of one probe per
tenant (packages/database/src/tenant-migrate.ts: one
information_schema.schemata lookup plus one information_schema.tables
lookup per sweep, parsed into two in-memory sets the per-tenant loop then
treats as authoritative).
Section 3 — Why sibling subaccounts share a schema (and what isolates them)
Subaccounts under one reseller parent share one tenant schema. Isolation between siblings therefore isn’t schema-vs-schema — it is org-ownership gating on every shared-schema surface. Each row in a shared-schema table carries an owning-org column, and every read or write on that table is gated against the requesting org:- Numbers, CNAM, and emergency addresses are gated by org ownership, exactly as the number lifecycle, CNAM, and emergency-address pages state: a sibling subaccount cannot read or mutate a record owned by another sibling, even though both live in the same schema.
- Wallet / transfer level: the parent moves capabilities between siblings (for example reassigning an owned number to a sibling), but only with ownership checks enforced on each verb.
- Public catalog rows (organizations, to resolve the tenant id; platform IP denylist; DNIS/inbound routing) are readable as catalog — not data — but each mutation is still ownership-checked.
Org-P with sibling subaccounts
Org-A and Org-B, all bound to one tenant schema tenant_p_9f8e7d6c:
The boundary to remember: sibling subaccounts share a schema; different
tenants never do. A sibling co-owns the schema; a different tenant has
its own schema entirely.
Section 4 — How the API resolves tenant on every request
The tenant is never taken from request input —tenant_id is a security
boundary, not a parameter. The resolution chain on every authenticated
request:
- Auth middleware resolves
sk_live_...to one org, reads that org’stenant_idandtenant_schema_readyfrompublic.organizations. - The per-request bridge opens the request against that org’s
tenant_<tenant_id>schema. If the schema is still provisioning and the route did not declare itself able to degrade, the API returns 503Tenant database temporarily unavailablehere — before any handler runs — rather than issue a query against the wrong schema. - The handler lists
sender_poolsfrom the resolved tenant schema and returns it. A client-suppliedtenant_idin the query string or body is ignored at this point: the resolved org from step 1 wins.
tenant_id values are ignored or rejected on the handful
of legacy endpoints that still accept them; the resolved org always wins.
Section 4.1 — How a connection reaches the tenant schema
The “per-request bridge” in step 2 is a real mechanism, not a request-time abstraction: a Fastify preHandler (apps/api/src/plugins/tenant.ts) calls createTenantClient(url, tenantId) (packages/database/src/tenant.ts) and, on failure, returns
the 503 above before the route handler runs. The handle it attaches is a
Drizzle client bound to one schema — and the way it binds is worth
stating plainly, because the obvious alternatives are not what runs:
- Not row-level security. Orbit does not gate tenant access with Postgres row-level security policies. The gate is schema-per-tenant (Sections 1–2): a request’s queries only resolve names through the resolved tenant’s search path, so there is no cross-tenant row to trip over in the first place.
- Not schema-qualified table names. Route handlers write plain,
unqualified Drizzle queries (
messages,sender_pools, …). The tenant binding is ambient, not per-relation. - Resolution: a transaction-scoped search path, not a per-tenant
session state. Each cached per-tenant pool passes
search_pathas a connection startup parameter, but — because the pooler runs in transaction mode and drops that parameter — the load-bearing enforcement is the wrapperwrapSqlWithTenantSearchPath(packages/database/src/tenant.ts): every emitted query runs inside a one-shotBEGIN; SET LOCAL search_path TO "tenant_<id>", public; <query>; COMMITblock. A commit re-scopes the backend before it returns to the pool, so no tenant path ever leaks onto a shared connection.
SET LOCAL (not bare SET) is the only
safe primitive here: a session-level SET search_path would leak
tenant state back into the pool’s backend connection for the next
request to inherit. Second, the , public suffix is deliberate: it is
what lets a tenant-scoped query also resolve the shared catalog tables
of Section 1 without qualification.
Section 4.2 — Connection scoping and pooling
A “tenant-scoped DB client” is, operationally, a cached postgres-js pool plus a schema-binding wrapper:- One pool per (pod, tenant), one connection per pool. Tenant work
is sequential at the API layer, so
createTenantClientcaps each tenant pool at a single underlying connection — the one-track queue is explicit, not implicit — and a second concurrent connection per (pod, tenant) buys no parallelism while inflating the cluster-wide connection count. Pools are cached per pod (LRU-evicted at theDEVOTEL_MAX_TENANT_POOLScap, default 512) and keyed by tenant, so a request against a previously-seen tenant reuses its pool — including its schema binding — rather than reconnecting. - Idle pools are reaped. An idle tenant pool closes after the
DEVOTEL_TENANT_POOL_IDLE_TIMEOUT_SECONDSwindow (the schedulers that touch many tenants on a slow cadence would otherwise hold one open per tenant forever), and the LRU cap bounds the worst-case connection footprint to the trailing working set. - No shared-credential isolation. There is no per-tenant database
role and no RLS: every connection authenticates with the platform role
and relies on the search-path gate above plus server-side org
resolution (Section 3’s ownership gating) — which is exactly why the
SOC 2 control in Section 5 frames
tenant_id-from-server as the load-bearing control.
Section 4.3 — The inbound reverse map: keeping tenant routing complete
Sections 1–4 describe how a request resolves tenant once the request knows which tenant it wants. Two surfaces resolve tenant the OTHER way around — from the destination number back to the owning tenant:- Inbound MO SMS / DLR webhooks arrive with no bearer token. The
resolver’s fast path looks up the destination E.164 in the shared
catalog table
public.phone_number_owners(one row per phone number, carryingtenant_id+org_id) — the O(1) reverse map. - A purchase or an admin SQL assign populates it best-effort; if
that write is refused (provisioning race, operator SQL, a mid-migration
window) the row may simply be missing — and the resolver’s own
fallback (a bounded per-tenant scan capped at
INBOUND_ORG_SCAN_CAP = 250orgs) will silently DROP the MO/DLR if the tenant sits beyond the cap.
processInboundReverseMapReconcile
(apps/webhook-worker/src/scheduler-inbound-reverse-map.ts), enumerated
independently of subscription status: it walks every schema-ready org via
getAllSchemaReadyTenantOrgs — deliberately WITHOUT the subscription
filter, because inbound resolution is not subscription-gated, so a tenant
whose Orb/Stripe subscription drops to cancelled/starter keeps it
schema AND its SMS-capable DIDs, and the same gap still resolves — and
backfills each org’s SMS-capable, owned-status DIDs into the reverse map
(idempotent INSERT ... ON CONFLICT DO NOTHING; a concurrent runtime
upsert wins).
Two gates keep the sweep defensible on a large fleet:
- Per-org failures gate through
isBenignReverseMapReconcileSkip— exactly the three benign/self-resolving classes (a mid-provisioning provisioning gap — the orphan-tenant-schema reaper’s concurrent drop; a pgbouncerquery timeoutbackstop; a SQLSTATE-less connection-class blip) are demoted to debug plus a counter, retried on the next hourly tick because the sweep is idempotent. A genuine fault still surfaces loudly — a non-provisioning-gap SQLSTATE (e.g.57014statement_timeout,53300too_many_connections— the sibling-retry-set survivor) matchesfindPgErrorNode(...).codeso its pg_code is never masked. - The tick’s self-check is the read-only probe
inbound-reverse-map-recon-sense.mjs— if it still sees a gap after the hourly tick, the probe alarms on the stuck org; a transient per-org skip never pages.
apps/api/src/scripts/reconcile-inbound-reverse-map.ts runs the
same backfill for a single tenant or as a --dry-run preview — the
scheduler is the steady-state write side; the CLI is for targeted / operator
runs.
Section 5 — The SOC 2 posture this underpins
The server-side resolution above is the load-bearing control the SOC 2 tenant-isolation controls page cites:tenant_id originates from the API key + the organizations catalog,
never from request input. That control, plus the schema-per-tenant data
split in Section 2 and the org-ownership gating in Section 3, is what the
auditable posture rests on.
Cross-references
- SOC 2 controls → Tenant Isolation — the
control text behind
tenant_id-from-server. - Messaging data residency — the tenant-owned, within-my-tenant region pin, and how it differs from the hard isolation boundary this page documents.
- Data placement and residency — the unified model that binds isolation (WHERE), the residency pin (WHERE-GEOGRAPHICALLY), and region-of-compute into one answer.
- Number lifecycle → Reassign to a subaccount — sibling-subaccount ownership gating in practice.
- SMPP guide — tenant-scoped
smpp_credentialsin action. - Scheduler fleet model — the bounded per-tenant tick pattern behind every repair/reverse-map/enumerate loop this page enumerates.