Skip to main content

The DLP scanner model

The DLP scanner feature page reads as settings and enforcement outcomes. This page is the model underneath it — how the scanner classifies what a candidate is, where on the outbound path the scan executes, what each channel hands the scanner, and what a custom rule would plug into if you ever author one.
The scanner is a tenant-owned control. Your organization owns the mode and the category set; the platform ships defaults and the pre-dispatch placement, and never flips your posture for you.

1. Pattern buckets, not a classifier monolith

DLP scanners in the wild group their detectors into pattern buckets by how a match is proven. Orbit’s DLP rule ships buckets you can reason about one at a time, and it deliberately stops short of machine-learned classification: The bucket decision is the scanner’s precision contract: every gate-grade bucket validates a candidate twice — shape, then checksum or registry — and the one class shape can’t pin (the bare passport token) is gated behind a dictionary anchor and kept off the default set. Where a bucket can’t reach gate-grade precision, the scanner doesn’t soften the rule; it excludes the class from the floor until you opt in. This is also what separates the DLP rule from the transcript scrubber: the storage-side scrubber accepts fuzzy matching because it redacts at write time, while the send-side gate must defend a hard refusal on a strict-mode send. Regex + validator buckets survive an audit; a probabilistic score on a send gate would not.

2. Pipeline placement — a pre-dispatch gate

The DLP scan does not run on its own queue or after dispatch. It is rule 7 in the policy scanner’s ordered rule chain, and the policy scanner itself runs as a single preHandler ahead of the send route’s handler — so a finding decides the request before the message is stored, queued, or dispatched:
  1. Your send request hits the route; the policy hook extracts channel, recipient, body, and metadata off the request.
  2. The ordered rules run on that payload; the DLP bucket runs on the body alongside TCPA quiet hours, SHAFT, spam score, and the sender gates.
  3. The DLP rule rolls into the same pass / warn / block verdict every other rule uses, and your organization’s enforcement mode decides what the verdict does at the transport.
Two placement facts make integration predictable: a body with no recipient or no content skips the scan cleanly, and a strict-mode refusal lands as POLICY_VIOLATION (HTTP 400) with nothing stored to reconcile. The same decision path is what the compose-time linter (POST /api/v1/messages/lint) calls synchronously, so the verdict your compose UI shows is the verdict the send will get. Fail-closed vs fail-open. The DLP rule itself is fail-open by design: it is pure and synchronous, has no external dependency to fail, and a body that matches nothing passes on shape alone. The fail-closed boundary sits one layer up — if the send hook cannot resolve your organization’s scan mode, the send is refused with a scanner-outage code rather than sent unscanned (covered in the pipeline model).

3. What each channel feeds the scanner

Every channel’s outbound body lands in the same scan, because a card number is sensitive on any lane — the differences are which fields of the envelope the scanner sees:
  • SMS — the text body. GSM/UCS2 encoding and segmentation don’t change detection; the scan runs on the composed body before segmentation.
  • WhatsApp — the body text (and a template’s filled parameters once rendered into the body).
  • Email — the body; the subject line is scanned by the sibling spam-keyword classifier, not the DLP categories.
  • RCS — the message text payload; a media-only message has no body to scan and skips cleanly.
The scanner reads the body field of the outgoing envelope; it does not open attachments or follow links (the sibling URL-reputation scan handles the link surface). Findings always come back as category + character offsets, never the matched text, on every channel alike.

4. Tenant-owned mode and category set

Two knobs decide what a finding does, and both belong to your organization. The organization enforcement mode (warn / strict / off) maps any verdict to a transport outcome — in warn a DLP block verdict still sends with the finding recorded; in strict the same verdict refuses the send. The DLP rule mode (block / redact / warn / off) decides what the finding itself is — whether it raises the verdict, strips the value into a [REDACTED_*] sentinel before dispatch, flags without gating, or doesn’t run. The interaction matrix and per-mode response shapes are on the feature page; the model takeaway is that category selection (which buckets run) and verdict plumbing (what a hit does) are independent dials on the same tenant-owned control.

5. Interaction with inbox AI privacy

The inbox AI privacy gates govern whether conversation content ships to a third-party LLM for categorization and summaries — an inbound, processor-scope decision. The DLP rule makes no LLM call at all: every bucket is deterministic string validation, so it scans exactly the same when your organization has turned the LLM hops off. A HIPAA-strict posture that disables the inbox AI gates therefore keeps its outbound DLP gate untouched — the two controls intersect only in that both draw a line around message content, one at the processor boundary, one at the dispatch boundary.

6. Authoring a scannable pattern

If you need a category beyond the four shipped buckets — an internal account identifier, a national ID format, a member number — the plug point is the dlp-scanner module in the compliance package. Choose the bucket first:
  1. Can a checksum or registry validate it? Write the candidate regex, then require the validator (Luhn-style, mod-97-style, an issuer range table) before the match counts — this is the only shape allowed to raise a block verdict on a send gate.
  2. Is shape alone ambiguous? Anchor it with a context-word dictionary gate, the same way the passport bucket requires a passport cue plus at least one digit in the token.
  3. Neither works? Keep it out of the send gate. A category that can only be fuzzy-matched belongs in a storage-side scrubber, not in a rule a strict-mode organization will hard-block on.
Whichever bucket you pick, keep the two invariants every shipped bucket holds: findings carry category + offsets only (a new author must never echo the matched value into a violation message, header, or audit event), and the scan stays pure and synchronous (no DB, no network — the linter runs it per keystroke). A change to the shipped categories goes through support today, as described on the feature page — treat new-bucket requests the same way.

7. Reading it with the rest of the map