Tenant provisioning lifecycle
The tenant isolation page explains what a tenant is — a catalog row inorganizations plus one PostgreSQL
schema named tenant_<tenant_id>, resolved per request. This page
explains how that tenant comes into being: the sequence that turns
an organization row into a schema-ready tenant, the two-phase split
that determines when you stop seeing 503, and the recovery queue a
stuck tenant enters. The individual failure codes the lifecycle emits
(MISSING_TENANT, TENANT_NOT_FOUND, TENANT_PROVISIONING, and
TENANT_SCHEMA_INCOMPLETE) are enumerated on the
troubleshooting ladder;
this is the concept page that ladder references.
The provisioning sequence
Provisioning is not one step — it is a fixed sequence, and the order is the whole story. A signup ( Clerk webhook ) or a sub-account create runs it:TENANT_PROVISIONING 503 rather than
double-creating. Steps 3→5 are the fast core path: the schema is
created synchronously and the ready flag flips within seconds, so
account-scoped APIs stop returning 503 within seconds of signup. Step 6
(the ~1,200-body TENANT_MIGRATIONS catch-up) is deliberately off the
request path because it gates nothing, and step 7 (the billing
customer) is likewise deferred — otherwise a slow provider call holds
signup open and a dropped webhook strands the user. Both are reconciled
by background sweeps, never by retrying the request.
Every statement in the schema bootstrap is CREATE ... IF NOT EXISTS,
inside one transaction with a transaction-scoped advisory lock — so a
concurrent second provisioning attempt for the same tenant blocks until
the winner commits, then sees the fully-built schema. A whole retry of
the bootstrap is therefore safe: nothing from a failed attempt survives
to conflict.
The titles above name what really matters: the flag
tenant_schema_ready is the one readiness signal every gate, scheduler
fan-out, and repair sweep keys off. A code like
TENANT_PROVISIONING (503) is emitted exactly when that flag is FALSE —
the lifecycle states on the
troubleshooting ladder
all derive from the same one flag.Where each lifecycle state comes from
Provisioning emits four bad responses on the ladder, and all four come from distinct points in the sequence above:MISSING_TENANT(400) — nothing in the request resolved to an org (no API key / tenant header); the sequence never even started.TENANT_NOT_FOUND(404) — the org row failed to land in step 2 (or the request resolves to a deleted/non-existent organization).TENANT_PROVISIONING(503) — the org row exists but steps 3–4 have not completed (flag FALSE), or have been reset FALSE by the repair loop below.TENANT_SCHEMA_INCOMPLETE(503) — the tenant is ready, but a later platform release added a table this tenant’s ledger has not yet applied; the literal 42P01undefined_tableSQLSTATE is what a schema-guarded route maps to this code.
When it goes wrong — the recovery queue
A stuck tenant is not repaired by the user refreshing the page. It enters one of three idempotent recovery paths, all of which re-run the sameensureTenantSchema bootstrap rather than a special repair:
- Per-request self-heal. The first request that finds
tenant_schema_ready = FALSEre-runs the bootstrap in-band, then the request resolves as if the sequence had completed. This covers the “flag stuck FALSE” case. - The 15-minute schema-repair sweep. A scheduler re-runs the bootstrap for an org whose flag is TRUE but whose schema is missing — the reverse gap no request will ever trigger. This covers the “schema was dropped or never fully landed” case.
- The deferred catch-up + periodic migration sweep. The background
runSingleTenantMigrationspass (step 6) reconciles any column an older body added; the scheduled fleet-widerunTenantMigrationssweep re-applies flagged per-tenant migrations so a long-lived tenant and a same-day signup converge.
_migrations
ledger says is missing, and one flag read is enough for the platform to
route the repair.
Relationship to isolation
Provisioning answers how the tenant boundary comes into being; tenant isolation then answers what that boundary is — the shared public catalog vs the per-tenant schema — and how the API resolves it per request — from your API key, never from a client-suppliedtenant_id. Read pages in that order; the
lifecycle states above are diagnosed on the
troubleshooting ladder
and the schema-drift-specific rung has its own runbook at
TENANT_SCHEMA_INCOMPLETE (503).
Cross-references
- Tenant isolation — what the boundary is and how resolution works.
- Troubleshoot tenant provisioning lifecycle — the four-code ladder and the fix per code.
- TENANT_SCHEMA_INCOMPLETE — the dedicated runbook for the drift code.
- Scheduler fleet model — the bounded per-tenant tick pattern behind the repair sweep above.