Scheduler fleet model
Orbit runs roughly 130 recurring background schedulers (tenant migrations,
reverse-map reconcile, ACD reconcilers, analytics rollups, agent batch,
among others). Every one of them is built on the same five building blocks.
This page names those blocks once so you can read any dashboard, scope any
dashboard-stale report, and interpret a 503/delayed-state question without
re-deriving the pattern from an individual scheduler’s code. The per-page
details live in tenant isolation (Section 2.1)
and operational sentinels; this page is
the shared model they each reference.
The five building blocks
1. Bounded enumeration over schema-ready tenants
Every scheduler tick starts by enumerating the tenants that are eligible
for work. The canonical enumeration helper, getAllTenantIds, reads the
organizations catalog and applies a three-part gate — active
subscription, not soft-deleted, and tenant_schema_ready = true — then
rejects malformed tenant identifiers against a strict v4-UUID shape so a
write-path leak can never carry a scheduler tick into a non-existent
schema. A wider variant (getAllSchemaReadyTenantIds /
getAllSchemaReadyTenantOrgs) drops the subscription-status gate for
sweeps whose work matters past the point of cancellation — number-rental
billing still matters, inbound resolution still matters — but always keeps
the schema-readiness gate.
The enumeration is deliberately bounded: a short-TTL, single-flight
cache collapses the ~130 schedulers sharing the same tenant list into a
single fleet-wide read, the result set is shuffled per accessor call so a
redeploy-heavy cluster does not starve the newest tenants of every sweep
(see the tail-starvation discussion on
tenant isolation),
and a stale-while-error fallback plus a re-query backoff stop the fleet
from stampeding Cloud SQL when the shared read misbehaves. Those bounds
are what make “one scheduler tick” a well-defined unit an operator can
reason about: it enumerates, it filters, it fans out.
2. Per-tenant tick classification
Within one tick, each tenant’s work is attempted individually and the
outcome is classified. The canonical classification union
(isBenignTenantSkipError, in create-tenant-scheduler.ts) recognizes
three non-fatal branches and demotes them to a quiet skip — no Sentry
event, no WARN log — while every genuine fault still surfaces loudly:
- Tenant timeout — a tenant whose previous tick is still in flight; the
scheduler skips this tick and lets the next tick finish it.
- Transient skip — a transient Postgres fault (connection refusal,
statement timeout under load, too-many-connections stampede,
serialization conflict, lock timeout; roughly the SQLSTATE classes
08*, 53300, 53400, 57P0*, 40001, 40P01, 57014, 55P03,
55006) matched by SQLSTATE so the skip gate cannot mask a real
defect.
- Degrade — a schema-shape fault a deliberately tolerant endpoint
elects to degrade on (a not-yet-migrated migration window on the
read/write guard), classified so a half-rolled-out migration never
converts into a wrong-tenant read.
Hand-rolled schedulers that do not use the shared factory still apply the
same three-branch classification in their per-tenant catch blocks — the
pattern is fleet-wide, not factory-specific.
3. Backoff and benign-skip naming discipline
Every scheduler exposes its cadence as a named constant (the 15-minute
tenant-schema-repair sweep, the hourly reverse-map reconcile, the 60-second
ACD pause timer, and so on), and its skip classifications use the
same vocabulary: transient_db_error, tenant_schema_missing,
degraded_*, skipped_*. That naming discipline is what lets a dashboard
or log line say “the tenant schema repair sweep skipped this tenant as
transient” and mean the same thing on every scheduler. Individual
schedulers are never permitted to invent their own names for the same
three branches.
4. Self-check probes
A recurring sweep is only useful if someone proves it still covers the
ground it claims. For each repair/reconcile loop, Orbit pairs the
write-side scheduler with a read-only probe that re-checks the same
invariant from the outside. The hourly inbound reverse-map reconcile pairs
with inbound-reverse-map-recon-sense.mjs; the tenant-schema repair loop
pairs with a read-only data-integrity sweep
(provisioning-tenant-schema-missing, critical severity). If the probe
still sees a gap after the tick should have healed it, the probe alarms;
if the tick skipped a tenant transiently, the probe stays quiet. This is
how the between-sweep drift window is kept finite instead of relying on
operator trust.
5. Code-pinned tick contracts
The tick contract long-lived enough to regress silently is locked in a
source-pin test (*.source-pin.test.ts / *.source-pin.mjs — plain
Node under apps/*/src or infrastructure/) that asserts the scheduler
still contains the classification branch, the readiness gate, the cadence
name, and the probe wiring. These pins are how each of the fleet’s
recurrence rules survives refactors, and they are the public, checkable
statement of what “a correct scheduler” means on this page.
Why you see this as 503s or delayed state
The building blocks above are precisely what generate the two customer-
visible behaviors operators ask about:
- 503 during provisioning or mid-migration. If a tenant is still
mid-
ensureTenantSchema, or reaches an endpoint whose route hasn’t
opted out of the ready gate, the schema-readiness filter in
getAllTenantIds and the request-layer attach in
tenant isolation
both convert it to a 503 rather than fire a query at the wrong schema.
- Delayed convergence between sweeps. A scheduler that wakes once an
hour can only converge drift once per tick; the self-check probe alarm
exists because, between ticks, drift is expected and baked in. An
elevated 503 rate with no corresponding probe alarm is a transient skip
the next tick will absorb, not a defect.
getAllTenantIds gates on active subscriptions and ready schemas; a “the
scheduler skipped me” check therefore means one of three deliberately
non-fatal branches above fired. A scheduler fault is the branch that can
never be silently skipped; benign branches skip precisely so they cannot
page.
See also
- Tenant isolation — the schema-ready flag
and the 503 generation boundary schedulers read for their fan-out.
- Operational sentinels — the settled-
but-non-outcome statuses a per-tenant tick classifies, so analytics
don’t count skips as failures.
- Number lifecycle — where the subscription-agnostic
schema-ready enumeration variant matters (rented DIDs per cancelled
tenant).