Skip to main content

Voice Guard Observability: the Deep-Dive Runbook

The voice-guard pages each cover one slice: voice guard documents the federal-window decision check, voice traffic-lane resolution documents the classifier upstream of it, and voice guard sentinel observability documents the breadcrumb contract. This page is the runnable runbook that joins them — how a tenant actually observes a guard, charts the event stream, alerts on its absence, and works the three failure modes to a resolution. The boundary stays where it always is: the federal TCPA voice window is the one control the platform owns (no toggle, fail-closed), and every other gate — your quiet-hours, lane claims, jurisdiction scope, any alert you build — is yours. This page documents tenant-owned observability only; it describes no platform-mandated gate beyond the federal floor and adds none.
Nothing here is legal advice. Which laws apply to your traffic depends on your jurisdiction, your recipients, and what you send — confirm with qualified counsel before you rely on this page’s framing.

1. Architecture: guard, classifier, sentinel

Three components produce the signal you observe, in a fixed order:
  1. The voice-lane classifier — the dispatcher stamps a lane (marketing or transactional) from your call metadata before any window math. The resolution precedence and the one-way ratchet (only traffic_lane: "marketing" is honored; a transactional claim is ignored) are on voice traffic-lane resolution. The classifier is the scope predicate: it decides whether the federal window even applies to this call.
  2. The TCPA federal guard — the 8 AM–9 PM recipient-local window check. It reads the resolved lane and the recipient timezone and returns one of five reason values (transactional_lane, non_us_recipient, inside_federal_window, outside_federal_window, timezone_unresolved). Evaluation chain and error surface: voice guard.
  3. The sentinel — a deliberately instrumented wrapper every gate runs through, never the raw guard directly. It emits one Sentry breadcrumb per invocation under the category compliance.tcpa.federal-window, fired before the guard evaluates, on allowed and blocked verdicts alike. The breadcrumb payload fields (recipientPrefix, mode, hasTimezoneOverride, trafficLane) are enumerated on sentinel observability.
Read the relationship this way: the classifier decides whether the window applies, the guard decides whether now is inside it, and the sentinel makes both questions answerable from a count rather than from code. Four dispatch surfaces run this chain — the API voice-compliance guard (ad-hoc/softphone), the campaign dialer pace, the callback dispatcher, and the voice-gateway pre-dial — and all four emit through the same sentinel category. The gate-by-gate precondition and blocked-outcome table is on sentinel breadcrumb troubleshooting.

2. Enabling and disabling layers — and the dialer preview

You cannot disable the federal guard; you can only layer your own controls on top, and you can choose how loudly the platform reports each evaluation. Three independent knobs, none of which weakens the federal floor:
  • Your tenant quiet-hours layer — opt-in, default-open, and the only layer you fully own. It can narrow within and below the federal window (your own hours, state overlays stricter than the floor) but never reopens a federally closed hour. The layering rule is worked out on state calling windows.
  • The evaluation mode at an integration point — enforce (default: a block denies the dial), warn (returns and logs the verdict, never throws — for dry-run pacing), or preview (returns the verdict silently — for scheduling previews). Mode changes how a block is reported to you; it never changes whether the guard ran, and the sentinel breadcrumb carries the mode in its mode field so your charts can separate dry-run signal from enforced signal.
  • The dialer preview check — the campaign preview surface runs the guard in preview mode so a launch screen can show “this contact would hold until X” without dispatching anything. Preview evaluations still emit their sentinel breadcrumb (the category is uniform across modes), which is the observability contract holding: a preview is a real evaluation, and its breadcrumb is how you confirm the preview path is wired. What preview suppresses is enforcement and the blocked-verdict service log — never the breadcrumb.
The consequence for your observability build: a tenant running a heavy preview workload will show a high breadcrumb volume in mode: preview with zero blocked dials, and that is healthy — previews are decoupled from holds by design. Filter your alerts on enforcement-carrying modes (enforce, warn) when you want a wiring signal on actual send paths, and treat preview volume as confirmation the planning surface is wired, not as dial traffic. The read-only planning endpoint semantics are on quiet-hours preview.

3. Charting the sentinel event

The sentinel’s value is that the wiring question reduces to a count. Build the chart on the breadcrumb category, grouped by gate and mode — never by message text, which differs across emission points. In Sentry (source of truth). The stream lives at breadcrumb.category:compliance.tcpa.federal-window. Chart it as an event count over time, split by mode and by the emitting service environment (API, dialer worker, callback dispatcher, voice gateway). The worked queries are on sentinel breadcrumb troubleshooting. Into your own Grafana or Datadog. The platform’s audit chain is tenant-exportable, so the same verdicts the sentinel marks can leave Orbit on a sink you own and land in your own dashboard:
  • Datadog — the native datadog sink streams the audit trail to your Datadog Logs intake; build a log-based metric on the guard-decision events and graph the count. Connector shape and activation calls: SIEM sinks.
  • Grafana — no first-party Grafana sink; route through the generic https_drain sink (a JSON POST per audit event) into a receiver your Grafana stack reads (for example Loki or a Prometheus-remote-write forwarder you operate), or pull on a schedule through the audit export API and chart offline. Either way the tenant owns the pipeline — the platform ships the events, your side stores and charts them.
Whichever sink you chart from, keep two series distinct: the verdict series (allowed vs blocked, by reason) answers “what is the guard deciding,” and the sentinel presence series (event count non-zero per gate) answers “is the guard wired.” The runbook’s alerting section below keys on the second.

4. Threshold: alert on a 30-day-zero rate

The one alert the sentinel is built for: if the sentinel event count for a gate stays at zero for more than 30 days on a send path you know is dialing, the wire-up never took. Thirty days is the margin that survives campaign pauses, seasonal quiet, and low-volume tenants; below it a zero count is ambiguous between “no traffic” and “no wiring.” Implement the alert against whichever chart you built in section 3:
  1. Aggregate the sentinel event count grouped by gate (by service environment, or by send path where you can distinguish them).
  2. Alarm when any group reads zero across a rolling 30-day window while that path’s dial volume is non-zero — the non-zero-dial qualifier is what separates “gate unwired” from “genuinely idle.”
  3. Scope out the legitimate zeros: the dialer-pace and callback gates only evaluate +1 NANP recipients, so a non-US-only or transactional-only tenant can show zero there healthily. Alert on the paths and recipient classes that should dial.
Tune the threshold down only after you have a baseline; a low-volume tenant can extend the window rather than lower the bar, because the failure the alert catches (a gate present in code but never reached in production) does not self-correct with time.

5. Troubleshooting the failure modes

Three failure shapes recur, and each has a distinct signature in the sentinel stream.

Gate reached, no event

The dial physically dispatched — the recipient’s phone rang or a 422 came back — but the sentinel stream shows nothing for that gate around the instant. This is the silent-wire-up drift the sentinel exists to catch: the guard is present in code, the gate fired (or should have), and the breadcrumb never emitted, so observability cannot confirm any of it. Because the breadcrumb fires before the evaluation, an absent breadcrumb is never “the evaluation happened but wasn’t logged” — it means the emission itself failed or the path bypassed the wrapper. Work it:
  1. Confirm the send path is one of the four instrumented gates and its preconditions were met (+1 NANP recipient for the pace/callback gates).
  2. Pull Sentry for the exact instant of the dispatch under the category; a block with no preceding breadcrumb is the escalate-now shape.
  3. Escalate as a platform defect with the send path and tenant context — dials are still enforced in this shape; only the verification fails.
The full “wiring didn’t take” triage, with worked Sentry queries, is on sentinel breadcrumb troubleshooting.

Gate unreached

The sentinel stream is healthy on other gates but one path shows nothing because the dial never reached the guard at all — not a silent sentinel, an upstream hold. Distinguish it from gate-reached-no-event by the missing dial itself: there is no dispatch, no 422, no ring. The usual causes are all upstream of the federal guard — a contact stuck in your tenant quiet-hours window, a DNC or STOP suppression, the FCC synthetic-voice consent gate, or a dialer-bypass-banned send path that never made it to a gate. Read the full send-time chain on send gates and the per-code fix map on dialing-window blocked calls before concluding the federal gate itself is unreachable.

Misconfigured jurisdiction set

The federal guard governs +1 NANP recipients; everything else skips the federal wire (non_us_recipient). A jurisdiction set is misconfigured when your traffic’s actual footprint and the scope you alert on disagree — you chart the federal sentinel for a tenant whose list is entirely non-NANP and see a healthy-looking zero, or you ignore the federal alert for a US-heavy tenant because “quiet hours cover it.” The reads:
  • A zero sentinel count on a non-US-only tenant is expected, not a wiring defect — scope the section-4 alert to the gates and recipient classes that should dial (this is the legitimate-zero carve-out).
  • A non-NANP recipient that should be US (a mis-set country code that reached the map by mistake) fails closed with timezone_unresolved — the guard ran, the breadcrumb exists, and the correction is the contact’s timezone, worked on the voice guard page.
  • Your own country-scope overlays are tenant-owned and layer under the federal floor; the consolidated federal-plus-overlays read is on state calling windows.

See also