Skip to main content

Usage alert rules on the Insights surface

The Usage alert rules page (/insights/usage-alert-rules, reachable from the Insights hub’s Usage alert rules tile) is the console for watching the transport and billing signals that move a CPaaS account. This is the walkthrough of that page: what each rule watches, how threshold and anomaly rules differ, what Evaluate now does, and how to read the fired-alerts feed. The full request/response contract for the backing API is a separate page — Usage & delivery anomaly alert rules.

1. What this page watches

This page watches transport and billing signals, not KPIs. The metric registry holds exactly three entries:
  • SMS delivery rate — the share of terminal outbound SMS that delivered.
  • Outbound message volume — the count of outbound messages sent.
  • Spend — wallet debits aggregated across the account.
These are operational tripwires: a delivery-rate collapse usually means a route or carrier is degrading, a volume surge usually means a compromised API key or SMS-pumping traffic, and a spend surge means bill shock in the making. If you are here to watch customer-experience or AI-cost numbers — CSAT, NPS, CES, LLM spend, AI containment — you want the KPI Alerts tile on the same hub (/insights/alerts), covered in KPI alerts on business metrics. The rule mechanics are the same on both surfaces; only the catalogue of metrics differs.

2. The signal set

Each rule targets one of the three metrics, and the evaluator samples it on two cadences — a continuous background sweep over every tenant that has at least one rule, and the manual tick behind Evaluate now: Because the sampling reuses the dashboard’s own filters and ledger, an alert fires on exactly the numbers the dashboard shows — there is no second counting convention to reconcile. Each rule then aggregates its metric two ways, depending on mode: across a trailing window you define in days (threshold), or against the newest complete UTC day compared with a learned baseline (anomaly).

3. Threshold rules

A threshold rule compares a static, fixed bound against the metric’s trailing-window value. In the dialog you pick a comparator (is below, is at or below, is above, is at or above) and a number on the metric’s native unit — a percent for delivery rate, a message count for volume, a USD figure for spend. The moment the windowed value crosses the bound, the rule notifies you in-app. Typical shapes: “delivery rate is below 95 over 7 days” guards a drop; “spend is above 100over1day"guardsasurge.Whenthemetricis∗∗Spend∗∗,thedialogoffersaone−click∗∗suggestedthrottle∗∗preset(‘isabove100 over 1 day" guards a surge. When the metric is **Spend**, the dialog offers a one-click **suggested throttle** preset (`is above 50/day over a 1-day window`) you can accept and edit instead of typing numbers. A threshold rule always carries both a comparator and a threshold — the create dialog and the API both reject one without the other.

4. Anomaly rules

An anomaly rule removes the number entirely. Instead of naming a bound, you flag a deviation versus the metric’s own learned baseline: the newest complete UTC day is scored against a learned 28-day history, and the rule fires when the day departs sharply from that shape — downward for delivery rate, upward for volume or spend. The suggested-throttle posture of this surface applies here in full: an anomaly rule alerts on deviation; it does not auto-rate-limit anything. No sending is throttled, no spend is capped, no route is changed — the rule’s only effect is the notification, and the action on it is yours. Pick anomaly mode when you do not know what “normal” is yet, or when a fixed number would need per-account tuning; pick threshold mode when the bound comes from an SLO you hold yourself. Anomaly mode takes no comparator or threshold — only the metric and the window.

5. Evaluate now

Evaluate now is the button in the page header. It runs every enabled rule immediately against the latest data — the same on-demand evaluation the API exposes as POST /usage/alert-rules/evaluate — so a scheduled sweep’s cadence never delays a check you want answered now. The button is disabled on a tenant with no rules. Each run refreshes the table’s Current reading and Status badge even when nothing breaches, so it doubles as a “check the signal right now” probe.

6. The fired-alerts feed

Below the rules table, the Fired alerts section lists the breach events that already landed — newest first, capped at the last 50. Each entry shows the rule name, a human-readable breach sentence (for example, “SMS delivery rate is 82.4%, below the 95% threshold over the last 7d.”), the fire time, and a View metric link that follows the event’s deep link into Insights. Two details matter in practice:
  • Cooldown. A rule emits an event only when its status transitions into breached, or when it stays in breach and the re-notify window (default 24 hours, configurable up to 30 days) has elapsed. A bad week does not re-page you every sweep.
  • Insufficient data never fires. A metric with no data in the window reports insufficient_data on the table rather than fabricating a 0% alert.
Notification destinations. Today the destination is the dashboard itself: a firing rule lands in the bell dropdown and the Notification Center with severity warning, titled Usage alert: <rule name>. Email or webhook fan-out is not on this surface yet — notify_channels accepts only in_app (GET /usage/alert-rules/events is the integration path if you want the alerts in your own systems). The triage, read, and dismiss semantics of the Notification Center are covered in In-dashboard notifications and the Notification Center. Deleting a rule does not erase history — fired alerts already sent are kept in the feed; the delete confirmation states it.

7. Who can view, who can manage

Reads are open: the page, the rules table, and the fired-alerts feed render for every member whose credentials carry the usage:read scope (member and viewer roles included). Managing — create, toggle, delete, Evaluate now — is restricted to org roles owner, admin, or developer, and creates, updates, and deletes land in the audit log with the acting user. This mirrors the parent Insights surface’s role model: the hub itself is open to read, management follows write-scoped roles.

8. Cost postures — rule or cap?

A usage alert rule is a notification posture; a usage cap is an enforcement posture. Route a billing anomaly into a cap decision when the signal demands more than an alert: a fixed daily spend bound with an automatic pause, a hard volume ceiling that stops sends outright. Usage & delivery anomaly alerts covers that billing-side detection frame — the alert that arrives when spend crosses the configured frame, and the point where you decide to harden it into a cap. Use this page’s rules as the early warning; use caps when the breach is unacceptable rather than merely notable.

See also