Skip to main content

Tenant isolation

Orbit is a multi-tenant platform. Every organization’s working data lives in its own PostgreSQL schema named tenant_<tenant_id> (dashes become underscores in the schema name). Shared catalog and routing tables — the things every tenant needs to read but no tenant owns — live in the public schema. This page defines the terms, catalogues both splits, and walks the request-resolution chain end to end so you can reconstruct it from the code.

What a tenant is

Three related terms appear across the docs:
  • Organization (org) — a row in the shared public.organizations catalog with its own id, tenant_id, and provisioning status (tenant_schema_ready). Your API key is issued against one org.
  • Tenant schema — one PostgreSQL schema (tenant_<tenant_id>) holding that org’s working data: contacts, messages, campaigns, agents, sender pools, and so on. Roughly 200 tables form this per-tenant catalog.
  • Subaccount — a separate org (row) under a reseller parent that shares the parent’s tenant schema. Sibling subaccounts can co-own one schema; ownership gating decides what each can see.
A bare org is a catalog entry; it becomes a tenant once provisioning creates its schema and the org’s tenant_schema_ready flag flips. On most endpoints a not-yet-ready schema returns 503; endpoints that explicitly tolerate provisioning-in-progress opt out.

Section 1 — The public schema: catalog tables every tenant reads

public holds tables every tenant reads, or tables that identify tenants themselves. Reading a row here doesn’t read another tenant’s data — the org boundary is still enforced by ownership columns. Mutations on a public table are gated the same way: the platform IP denylist, the DNIS routing table, and every per-org row in organizations all carry an org-ownership check on write. public is catalog-only: no per-tenant working data (messages, contacts, recordings) ever lives there. These are catalog or routing tables: the shape is “one row per org or one row per platform artifact”, never “one table per tenant”.

Section 2 — Tenant schemas: one per tenant, named tenant_<tenant_id>

Everything an org creates, sends, or receives lives in its own tenant schema. Senders, inbound SMS routing, SMPP credentials, sender pools, contacts, campaigns, flows, agent state, recordings, inbox, CDP events, WFM — all are per-tenant tables. Provisioning creates the schema, builds the full base table set synchronously, and flips tenant_schema_ready; the tenant migration set then reconciles any drift in the background (Section 2.1). Because each tenant schema carries its own copy of the same table set, a pool id — or any other per-tenant resource id — is only unique within one tenant schema. sender_pools ids repeat across tenants; Postgres enforces uniqueness per-tenant via the table’s primary key inside each schema. A sibling subaccount shares your tenant schema, so a number it owns is visible to the full org-family tree; a sender pool it owns is visible too, but only inside your family’s schema — never across tenants.

Section 2.1 — How tenant schemas are provisioned and migrated

A schema is not just CREATE SCHEMA: it must carry the ~200-table set, stay consistent with every other tenant’s copy of that set, and report its readiness through one flag (organizations.tenant_schema_ready). The mechanism is split in two so the customer-facing path never waits on the slow path. Provisioning is a two-phase split — fast core, deferred catch-up. Signup and every other provisioning path run ensureTenantSchema(url, tenantId) (packages/database/src/tenant.ts) synchronously: inside one transaction it CREATE SCHEMA IF NOT EXISTS, SETs search_path, and builds the full canonical table set — every base table a customer-facing API touches exists the moment it returns — and the provisioning path flips tenant_schema_ready = true. That is a fast-core path: account-scoped APIs stop returning 503 within seconds of signup instead of waiting for the full migration walk. The slower catch-up — runSingleTenantMigrations(tenantId) (packages/database/src/tenant-migrate.ts), dispatched fire-and-forget by apps/api/src/lib/provision-tenant-migrations.ts — is an idempotent no-op on a fresh schema: every migration body is guarded by IF NOT EXISTS / to_regclass, so re-running it reconciles any column an older migration added without ever blocking the request. Migrations are versioned per tenant and drift-repaired on three tracks. The tenant migration set (TENANT_MIGRATIONS in packages/database/src/tenant-migrate.ts) is one composed array — every public-schema migration in the repo plus every tenant-schema migration — dispatched one body at a time; only bodies that reach the tenant schema run per-tenant (Section 1’s catalog tables are built there by the public-schema bodies in the same set). Each tenant schema carries a _migrations ledger recording which numbered migration ids have been applied to that schema. New platform releases append new numbered tenant migrations; the periodic runTenantMigrations() sweep (packages/database/src/tenant-migrate.ts) enumerates every tenant and applies only the ones each schema’s ledger has not recorded, so a long-lived tenant and a same-day signup converge to the same shape without the sweep ever re-running already-applied bodies. Migration order is fixed by the array — arrivals converge to release order, never to per-tenant signup order. Because the sweep is periodic, a newly released migration always has an until-next-sweep window in which some tenant schemas genuinely lack its table or column; the read-path and write-path guards below exist to cover exactly that window, and every schema change has to be tolerant of it during rollout.
  • Read-path: withTenantSchemaGuard (apps/api/src/lib/with-tenant-schema-guard.ts, Section 4’s attach path) converts a missing-table/missing-column error on a not-yet-migrated tenant into a structured TENANT_SCHEMA_INCOMPLETE 503 (or an empty degrade on read-only endpoints) rather than a wrong-tenant read — and the auth middleware’s healTenantSchemaIfNeeded (packages/auth/src/middleware.ts) re-runs the same idempotent ensureTenantSchema in-band on the first request that finds an org stuck at tenant_schema_ready = false.
  • Write-path: the 15-minute scheduler-tenant-schema-repair (apps/webhook-worker/src/scheduler-tenant-schema-repair.ts) sweeps for the reverse drift — an org whose flag is already true but whose schema is missing — and re-provisions it. The per-request self-heal only fires when the flag is false, so this scheduler is the repair loop for the “flag ready, schema gone” case no request ever triggers.
  • Reconciliation sweep: nobody has to verify a background apply dropped a tenant. runTenantMigrationsInBackground’s own failure never rolls the org back: the periodic runTenantMigrations() sweep and the per-request self-heal both re-run the same idempotent pass, so a dropped or partial apply converges on the next sweep with no manual step.
tenant_schema_ready and the 503 boundary. Until provisioning completes, most endpoints return 503 instead of querying a half-built schema; the flag is the single readiness signal every gate (Section 4), scheduler fan-out, and repair loop keys off. Two enumerations keep that cheap at scale: scheduler fan-outs read the ready flag from organizations through the short-TTL-cached, single-flight getAllTenantIds (apps/webhook-worker/src/lib/get-all-tenant-ids.ts) rather than probing each tenant schema, and the runTenantMigrations() sweep resolves every candidate schema’s existence — and which schemas carry a _migrations table — in two batched information_schema queries for the whole fleet, O(1) round trips instead of one probe per tenant (packages/database/src/tenant-migrate.ts: one information_schema.schemata lookup plus one information_schema.tables lookup per sweep, parsed into two in-memory sets the per-tenant loop then treats as authoritative).

Section 3 — Why sibling subaccounts share a schema (and what isolates them)

Subaccounts under one reseller parent share one tenant schema. Isolation between siblings therefore isn’t schema-vs-schema — it is org-ownership gating on every shared-schema surface. Each row in a shared-schema table carries an owning-org column, and every read or write on that table is gated against the requesting org:
  • Numbers, CNAM, and emergency addresses are gated by org ownership, exactly as the number lifecycle, CNAM, and emergency-address pages state: a sibling subaccount cannot read or mutate a record owned by another sibling, even though both live in the same schema.
  • Wallet / transfer level: the parent moves capabilities between siblings (for example reassigning an owned number to a sibling), but only with ownership checks enforced on each verb.
  • Public catalog rows (organizations, to resolve the tenant id; platform IP denylist; DNIS/inbound routing) are readable as catalog — not data — but each mutation is still ownership-checked.
Worked illustration — a reseller parent Org-P with sibling subaccounts Org-A and Org-B, all bound to one tenant schema tenant_p_9f8e7d6c: The boundary to remember: sibling subaccounts share a schema; different tenants never do. A sibling co-owns the schema; a different tenant has its own schema entirely.
Cross-tenant isolation exists at the schema layer; sibling-subaccount isolation exists at the org-ownership layer. The two are enforced differently — do not assume “same schema” means “open to the whole reseller family.” Each endpoint carries an ownership check.

Section 4 — How the API resolves tenant on every request

The tenant is never taken from request input — tenant_id is a security boundary, not a parameter. The resolution chain on every authenticated request:
Worked resolution, GET /v1/sender-pools:
  1. Auth middleware resolves sk_live_... to one org, reads that org’s tenant_id and tenant_schema_ready from public.organizations.
  2. The per-request bridge opens the request against that org’s tenant_<tenant_id> schema. If the schema is still provisioning and the route did not declare itself able to degrade, the API returns 503 Tenant database temporarily unavailable here — before any handler runs — rather than issue a query against the wrong schema.
  3. The handler lists sender_pools from the resolved tenant schema and returns it. A client-supplied tenant_id in the query string or body is ignored at this point: the resolved org from step 1 wins.
Client-supplied tenant_id values are ignored or rejected on the handful of legacy endpoints that still accept them; the resolved org always wins.

Section 4.1 — How a connection reaches the tenant schema

The “per-request bridge” in step 2 is a real mechanism, not a request-time abstraction: a Fastify preHandler (apps/api/src/plugins/tenant.ts) calls createTenantClient(url, tenantId) (packages/database/src/tenant.ts) and, on failure, returns the 503 above before the route handler runs. The handle it attaches is a Drizzle client bound to one schema — and the way it binds is worth stating plainly, because the obvious alternatives are not what runs:
  • Not row-level security. Orbit does not gate tenant access with Postgres row-level security policies. The gate is schema-per-tenant (Sections 1–2): a request’s queries only resolve names through the resolved tenant’s search path, so there is no cross-tenant row to trip over in the first place.
  • Not schema-qualified table names. Route handlers write plain, unqualified Drizzle queries (messages, sender_pools, …). The tenant binding is ambient, not per-relation.
  • Resolution: a transaction-scoped search path, not a per-tenant session state. Each cached per-tenant pool passes search_path as a connection startup parameter, but — because the pooler runs in transaction mode and drops that parameter — the load-bearing enforcement is the wrapper wrapSqlWithTenantSearchPath (packages/database/src/tenant.ts): every emitted query runs inside a one-shot BEGIN; SET LOCAL search_path TO "tenant_<id>", public; <query>; COMMIT block. A commit re-scopes the backend before it returns to the pool, so no tenant path ever leaks onto a shared connection.
Two properties follow. First, SET LOCAL (not bare SET) is the only safe primitive here: a session-level SET search_path would leak tenant state back into the pool’s backend connection for the next request to inherit. Second, the , public suffix is deliberate: it is what lets a tenant-scoped query also resolve the shared catalog tables of Section 1 without qualification.

Section 4.2 — Connection scoping and pooling

A “tenant-scoped DB client” is, operationally, a cached postgres-js pool plus a schema-binding wrapper:
  • One pool per (pod, tenant), one connection per pool. Tenant work is sequential at the API layer, so createTenantClient caps each tenant pool at a single underlying connection — the one-track queue is explicit, not implicit — and a second concurrent connection per (pod, tenant) buys no parallelism while inflating the cluster-wide connection count. Pools are cached per pod (LRU-evicted at the DEVOTEL_MAX_TENANT_POOLS cap, default 512) and keyed by tenant, so a request against a previously-seen tenant reuses its pool — including its schema binding — rather than reconnecting.
  • Idle pools are reaped. An idle tenant pool closes after the DEVOTEL_TENANT_POOL_IDLE_TIMEOUT_SECONDS window (the schedulers that touch many tenants on a slow cadence would otherwise hold one open per tenant forever), and the LRU cap bounds the worst-case connection footprint to the trailing working set.
  • No shared-credential isolation. There is no per-tenant database role and no RLS: every connection authenticates with the platform role and relies on the search-path gate above plus server-side org resolution (Section 3’s ownership gating) — which is exactly why the SOC 2 control in Section 5 frames tenant_id-from-server as the load-bearing control.

Section 4.3 — The inbound reverse map: keeping tenant routing complete

Sections 1–4 describe how a request resolves tenant once the request knows which tenant it wants. Two surfaces resolve tenant the OTHER way around — from the destination number back to the owning tenant:
  • Inbound MO SMS / DLR webhooks arrive with no bearer token. The resolver’s fast path looks up the destination E.164 in the shared catalog table public.phone_number_owners (one row per phone number, carrying tenant_id + org_id) — the O(1) reverse map.
  • A purchase or an admin SQL assign populates it best-effort; if that write is refused (provisioning race, operator SQL, a mid-migration window) the row may simply be missing — and the resolver’s own fallback (a bounded per-tenant scan capped at INBOUND_ORG_SCAN_CAP = 250 orgs) will silently DROP the MO/DLR if the tenant sits beyond the cap.
The write-side counterpart is the hourly scheduler processInboundReverseMapReconcile (apps/webhook-worker/src/scheduler-inbound-reverse-map.ts), enumerated independently of subscription status: it walks every schema-ready org via getAllSchemaReadyTenantOrgs — deliberately WITHOUT the subscription filter, because inbound resolution is not subscription-gated, so a tenant whose Orb/Stripe subscription drops to cancelled/starter keeps it schema AND its SMS-capable DIDs, and the same gap still resolves — and backfills each org’s SMS-capable, owned-status DIDs into the reverse map (idempotent INSERT ... ON CONFLICT DO NOTHING; a concurrent runtime upsert wins). Two gates keep the sweep defensible on a large fleet:
  • Per-org failures gate through isBenignReverseMapReconcileSkip — exactly the three benign/self-resolving classes (a mid-provisioning provisioning gap — the orphan-tenant-schema reaper’s concurrent drop; a pgbouncer query timeout backstop; a SQLSTATE-less connection-class blip) are demoted to debug plus a counter, retried on the next hourly tick because the sweep is idempotent. A genuine fault still surfaces loudly — a non-provisioning-gap SQLSTATE (e.g. 57014 statement_timeout, 53300 too_many_connections — the sibling-retry-set survivor) matches findPgErrorNode(...).code so its pg_code is never masked.
  • The tick’s self-check is the read-only probe inbound-reverse-map-recon-sense.mjs — if it still sees a gap after the hourly tick, the probe alarms on the stuck org; a transient per-org skip never pages.
The CLI apps/api/src/scripts/reconcile-inbound-reverse-map.ts runs the same backfill for a single tenant or as a --dry-run preview — the scheduler is the steady-state write side; the CLI is for targeted / operator runs.

Section 5 — The SOC 2 posture this underpins

The server-side resolution above is the load-bearing control the SOC 2 tenant-isolation controls page cites: tenant_id originates from the API key + the organizations catalog, never from request input. That control, plus the schema-per-tenant data split in Section 2 and the org-ownership gating in Section 3, is what the auditable posture rests on.

Cross-references