Skip to main content

Tenant schema migration and evolution

A tenant schema is not static after provisioning. As Devotel Orbit ships new features, the schema for each organization must converge on the shape expected by the current release. This page explains that evolution contract: what a tenant migration changes, how the platform retries safely, and what happens when a pass cannot finish. For the one-time sequence that creates an organization and its first schema, see Tenant provisioning lifecycle. For the boundary that keeps each organization’s data separate, see Tenant isolation.

What a tenant schema migration is

A tenant schema migration is a versioned, idempotent pass applied to a tenant_<tenant_id> PostgreSQL schema after provisioning. It adds or adjusts the tables and columns that a later platform release expects. The tenant migration set is distinct from the public-schema migration set: a tenant migration changes one organization’s schema and must never move that organization’s working data into public or another tenant. The schema bootstrap creates the current core table set quickly enough for signup to return. The migration pass then reconciles migration-added details that older schemas, or an older bootstrap shape, may not have. This is why a long-lived organization and a newly provisioned organization can start from different release histories but still converge on the same schema shape.

The three convergence paths

Orbit deliberately has more than one way to close the gap between a release and a tenant schema. Each path runs the same migration definitions and is safe to overlap with the others.

1. Provisioning background pass

After ensureTenantSchema() builds the core schema, provisioning starts runSingleTenantMigrations(tenantId) from apps/api/src/lib/provision-tenant-migrations.ts. The call is fire-and-forget: it runs off the signup hot path so the tenant_schema_ready flag can become usable without waiting for a full migration walk. This pass is especially important for a new organization. It catches columns that an older migration added but the base bootstrap may not yet mirror. A failure does not make an already-live organization unready; the other paths below get another chance.

2. Periodic tenant sweep

The periodic runTenantMigrations() sweep in packages/database/src/tenant-migrate.ts enumerates tenant schemas and applies only the migrations each schema still needs. It is the fleet-wide convergence path for organizations that were provisioned before a release, missed a background attempt, or were temporarily unavailable during one. The pre-deploy apps/api/src/scripts/run-tenant-migrations.ts job can front-load this same sweep for a release. API startup and scheduled passes remain safety nets, so a partial run resumes rather than requiring a manual schema reconstruction.

3. Request-path self-heal

The auth middleware’s healTenantSchemaIfNeeded runs when a request finds an organization whose schema readiness flag is still false. It re-runs the idempotent schema bootstrap in-band, then lets the request proceed when the heal succeeds. A lock held by another replica or a failed heal produces the normal provisioning response; it does not cause two replicas to build the schema independently. This path repairs provisioning-state drift. The periodic sweep and the background pass handle migration catch-up after the schema is live. Together, the three paths cover the initial signup, fleet-wide release evolution, and a request that discovers a provisioning gap.

Idempotent by construction

Every tenant migration body is written so that applying it again is a no-op once the schema has the requested shape. In practice, bodies use guards such as IF NOT EXISTS and to_regclass before creating or altering tenant-owned objects. The provisioner documents this contract directly in apps/api/src/lib/provision-tenant-migrations.ts. Idempotence matters because the paths can overlap: a provisioning pass can race the periodic sweep, and a retry can follow a transient database error. The migration runner also takes per-migration advisory locks and rechecks the ledger while holding the lock before executing a body. A retry therefore converges the schema instead of applying the same change twice.

The per-tenant _migrations ledger

Each tenant schema has its own _migrations table. It records a migration’s stable identifier and the time it was applied. The ledger is part of the schema boundary: tenant_acme has its own applied set, independent of tenant_other. When a runner considers a migration, it checks that tenant’s ledger. After the migration body completes successfully, it records the identifier in the same transaction. On a later pass, recorded identifiers are skipped, while an unrecorded identifier is applied and then recorded. This is why the sweep can revisit every tenant without replaying already-applied work. The ledger is not a claim that every guarded statement changed a row. A migration may legitimately find that its target table is absent and do nothing; the important contract is that the body is safe to revisit and that the tenant follows the ordered migration set. The daily drift reconciler covers the historical case where a guarded column was recorded but never landed.

Failure posture: live organizations stay live

A background migration failure never rolls back the organization. The schema is already live by the time the deferred catch-up starts, so the failure is logged, captured by Sentry, and counted as provisioning.background_migrations_failed. The platform leaves the pending work for runTenantMigrations() and the request-path self-heal to retry. This separation prevents a transient database or lock problem from turning a working organization into a permanently stuck 503. A release may leave a short convergence window while a tenant catches up, but the next safe path reuses the same guarded migration rather than requiring destructive rollback.

The daily column-drift safety net

apps/api/src/jobs/tenant-column-drift-reconcile.cron.ts is the daily safety net for a specific class of drift: a guarded ADD COLUMN can complete without adding anything when its target table was not present, yet the migration row could still look applied. The reconciler checks the catalog and the tenant’s own _migrations ledger, identifies confirmed missing columns, and hands the migration back to the canonical runner by removing its tracking row when repair is enabled. It does not become a second migration engine and does not run ad hoc DDL. The canonical runner re-applies the real, idempotent migration with its normal checks. This keeps one source of schema truth while giving historical drift a way back into the normal convergence loop.

Relationship to the sibling concepts

  • Tenant provisioning lifecycle owns the one-time organization-to-schema arc. Its deferred catch-up step is the first place to see the background runSingleTenantMigrations pass.
  • Tenant isolation owns the schema boundary and request-time tenant resolution. Every migration described here must stay inside the resolved tenant schema; public catalog tables remain a separate concern.
  • Tenant schema incomplete troubleshooting covers the response to a schema that has not yet converged. Use it when a missing table or column appears at a request boundary.

What this page is not

This is a concept model, not a database-upgrade runbook. You do not need to run SQL manually, edit the per-tenant ledger, or replay migration bodies from an operator shell. It is also not a guide to public-schema migrations. Those follow a separate ownership and rollout boundary; this page covers the per-tenant evolution contract only.