Analytics pipeline: from raw events to an aggregate read surface
Orbit aggregates every message, call, and spend event into the datasets that back the dashboard’s Insights pages and the public Analytics API. This page covers how that path is built — what the read surface exposes, why every analytics query targets the read replica, the rate limits that apply, the background producers that pre-aggregate the series, and how the dashboard consumes the same numbers. Endpoint-by-endpoint request shapes live on the Analytics API recipes page; this page is the model behind them.What the analytics surface exposes
Every analytics dataset lives under/api/v1/analytics and returns counts
over a caller-chosen window:
- Delivery metrics — sent / delivered / failed / read totals for message channels, plus a time-series of the same counters.
- Voice metrics — answered / completed / failed call totals over the same window.
- Cost — billed message count, spend per currency, average cost per billed message, and SMS segment counts.
- Breakdowns — the same counters split by channel, by destination country, and by error code for failed sends.
group_bybuckets the time-series athour,day,week, ormonthresolution. Day buckets are the default.- Filters narrow the window:
channel,status,campaign_id, and astart_date/end_daterange (or a rollingdayslookback). Both default to the last 30 days and are capped at 365.
analytics:read;
dashboard sessions pass through the route’s session authentication. A raw
API key that holds only send scopes is rejected with 403 before any query
runs.
Every response uses the standard Orbit envelope: the payload under data,
request_id and timestamp under meta.
The read-replica separation invariant
All analytics queries execute against the read replica, never the primary database. That separation is deliberate: aggregate reporting over long windows scans or joins large portions of the per-tenant message tables, and it would become contention on the same database that serves the live send path — the queue that posts outbound sends, the webhooks that mark deliveries, the dialer that owns active calls. The invariant in one sentence: reporting load never competes with live traffic for the primary database. It holds for the API and the dashboard alike — Insights reads are routed the same way as external API reads. One consequence for readers: the replica lags the primary by a small replication delta, so a metric you read reflects the fleet a moment ago, not the nanosecond a send finished. For analytics workloads that delta is immaterial; a minute-level aggregate sees the same answer.Rate limits
Analytics endpoints are rate-limited at 60 requests per minute per tenant, with the limit keyed on the organization. The cap exists because aggregate queries are heavier than CRUD reads even against the replica — they bound the burst capacity of reporting workloads without touching the per-endpoint send quotas on the live channels. A response over the cap returns429 with a Retry-After header. Honor
it — the limiter is per tenant, so your own dashboards and scheduled
exports see the same meter in the account.
How the roll-ups pre-aggregate the series
The hour/day/week/month buckets the API returns are not always computed on-demand. A pair of background producers in the webhook worker pre-aggregate message counters into per-hour buckets so the dashboards and the most common queries read a pre-built, indexed row instead of scanning raw events at request time:-
scheduler-stats-rollup.ts— aggregates the previous full hour of the tenant’s message table into a keyed(hour_utc, channel, status)row, with the count and the spend. The upsert is idempotent (INSERT ... ON CONFLICT DO UPDATE), so a late delivery webhook that backfills a closed hour re-aggregates cleanly instead of duplicating. Tenants are visited in small parallel batches and the pass returns early on shutdown; a skipped tenant re-rolls the same hour on the next tick. -
scheduler-analytics.ts— a broader producer that fans the same kind of per-tenant aggregation across additional analytics series: channel-engagement combined with sentiment and queue-quality roll-ups, so Insights panels that mix delivery with engagement read one aggregate instead of recomputing joins on raw rows.
How the dashboard consumes these datasets
The Insights pages in the dashboard read the same/api/v1/analytics
datasets the public API serves — the charts showing outbound volume by
channel, delivery-rate trend lines, per-country breakdowns, and spend per
currency are rendered straight over those endpoints. The scheduled-reports
surface a tenant configures in the dashboard composes the same
aggregates into a recurring e-mailed export.
That shared read surface is the point of the pipeline: the same
pre-aggregated rows power the dashboard view a user opens, the BI pull a
customer schedules, and the canned scheduled report — so a figure a
stakeholder sees on a dashboard is the figure an API script can fetch,
counted the same way.