> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orbit.devotel.io/llms.txt
> Use this file to discover all available pages before exploring further.

# The DLP scanner model

> How the outbound DLP scanner is organized as pattern buckets (validated regexes, keyword dictionaries, and scorers rather than ML classifiers), where the scan sits on the send pipeline, when it fails closed vs fail-open, what each channel contributes to the scanned text, and how to reason about writing a custom rule.

# The DLP scanner model

The [DLP scanner feature page](/compliance/dlp-scanner) reads as settings and
enforcement outcomes. This page is the model underneath it — how the scanner
classifies what a candidate *is*, where on the outbound path the scan
executes, what each channel hands the scanner, and what a custom rule would
plug into if you ever author one.

<Note>
  The scanner is a **tenant-owned control**. Your organization owns the mode
  and the category set; the platform ships defaults and the pre-dispatch
  placement, and never flips your posture for you.
</Note>

***

## 1. Pattern buckets, not a classifier monolith

DLP scanners in the wild group their detectors into **pattern buckets** by
how a match is proven. Orbit's DLP rule ships buckets you can reason about
one at a time, and it deliberately stops short of machine-learned
classification:

| Bucket | How a hit is proven | Ships as |
| - | - | - |
| **Regex-only** | Shape alone. Cheap and fast, but shape alone can't tell a card number from an order id, so no shipped category is regex-only. | Never a gate by itself — always paired with validation below. |
| **Regex + deterministic validator** | Shape first, then a checksum or registry table: Luhn mod-10 plus the card-network prefix/length table for a PAN; SSA never-assigned ranges for an SSN; ISO 7064 mod-97 plus the per-country IBAN length registry. | `credit_card`, `ssn`, `iban` — the default universal floor. |
| **Keyword/dictionary gate** | A token matches only when a context word from a fixed list appears next to it. Precision comes from the anchor word, not the token shape. | `passport` (opt-in) — the token fires only when it directly follows a `passport` context cue and carries a digit. |
| **ML classifier** | A scored model over features. The DLP rule ships **none** — a `block` verdict that a strict-mode organization turns into a refused send cannot rest on a probability the operator can't audit. | Not used on the send gate. The sibling [spam-keyword scorer](/compliance/policy-scanner) is the closest shipped scorer, and it is a deterministic points table, not a model. |

The bucket decision is the scanner's precision contract: **every gate-grade
bucket validates a candidate twice** — shape, then checksum or registry — and
the one class shape can't pin (the bare passport token) is gated behind a
dictionary anchor *and* kept off the default set. Where a bucket can't reach
gate-grade precision, the scanner doesn't soften the rule; it excludes the
class from the floor until you opt in.

This is also what separates the DLP rule from the transcript scrubber: the
storage-side scrubber accepts fuzzy matching because it redacts at write
time, while the send-side gate must defend a hard refusal on a strict-mode
send. Regex + validator buckets survive an audit; a probabilistic score on a
send gate would not.

***

## 2. Pipeline placement — a pre-dispatch gate

The DLP scan does not run on its own queue or after dispatch. It is **rule 7
in the policy scanner's ordered rule chain**, and the policy scanner itself
runs as a single `preHandler` ahead of the send route's handler — so a
finding decides the request **before the message is stored, queued, or
dispatched**:

1. Your send request hits the route; the policy hook extracts channel,
   recipient, body, and metadata off the request.
2. The ordered rules run on that payload; the DLP bucket runs on the body
   alongside TCPA quiet hours, SHAFT, spam score, and the sender gates.
3. The DLP rule rolls into the same `pass` / `warn` / `block` verdict every
   other rule uses, and your organization's enforcement mode decides what the
   verdict does at the transport.

Two placement facts make integration predictable: a body with no recipient or
no content skips the scan cleanly, and a `strict`-mode refusal lands as
`POLICY_VIOLATION` (HTTP 400) with nothing stored to reconcile. The same
decision path is what the compose-time linter (`POST /api/v1/messages/lint`)
calls synchronously, so the verdict your compose UI shows is the verdict the
send will get.

**Fail-closed vs fail-open.** The DLP rule itself is fail-open by design: it
is pure and synchronous, has no external dependency to fail, and a body that
matches nothing passes on shape alone. The **fail-closed** boundary sits one
layer up — if the send hook cannot resolve your organization's scan mode, the
send is refused with a scanner-outage code rather than sent unscanned
(covered in the [pipeline model](/concepts/policy-scan-pipeline-model)).

***

## 3. What each channel feeds the scanner

Every channel's outbound body lands in the same scan, because a card number
is sensitive on any lane — the differences are which fields of the envelope
the scanner sees:

* **SMS** — the text body. GSM/UCS2 encoding and segmentation don't change
  detection; the scan runs on the composed body before segmentation.
* **WhatsApp** — the body text (and a template's filled parameters once
  rendered into the body).
* **Email** — the body; the subject line is scanned by the sibling
  spam-keyword classifier, not the DLP categories.
* **RCS** — the message text payload; a media-only message has no body to
  scan and skips cleanly.

The scanner reads the body field of the outgoing envelope; it does not open
attachments or follow links (the sibling
[URL-reputation scan](/concepts/url-reputation-smishing-scan) handles the
link surface). Findings always come back as category + character offsets,
never the matched text, on every channel alike.

***

## 4. Tenant-owned mode and category set

Two knobs decide what a finding does, and both belong to your organization.
The organization **enforcement mode** (`warn` / `strict` / `off`) maps any
verdict to a transport outcome — in `warn` a DLP `block` verdict still sends
with the finding recorded; in `strict` the same verdict refuses the send. The
**DLP rule mode** (`block` / `redact` / `warn` / `off`) decides what the
finding itself is — whether it raises the verdict, strips the value into a
`[REDACTED_*]` sentinel before dispatch, flags without gating, or doesn't
run. The interaction matrix and per-mode response shapes are on the
[feature page](/compliance/dlp-scanner#choosing-a-polarity-and-what-comes-back);
the model takeaway is that category selection (which buckets run) and verdict
plumbing (what a hit does) are independent dials on the same tenant-owned
control.

***

## 5. Interaction with inbox AI privacy

The [inbox AI privacy gates](/compliance/inbox-ai-privacy) govern whether
conversation content ships to a third-party LLM for categorization and
summaries — an inbound, processor-scope decision. The DLP rule makes **no LLM
call at all**: every bucket is deterministic string validation, so it scans
exactly the same when your organization has turned the LLM hops off. A
HIPAA-strict posture that disables the inbox AI gates therefore keeps its
outbound DLP gate untouched — the two controls intersect only in that both
draw a line around message content, one at the processor boundary, one at
the dispatch boundary.

***

## 6. Authoring a scannable pattern

If you need a category beyond the four shipped buckets — an internal account
identifier, a national ID format, a member number — the plug point is the
[`dlp-scanner` module](/compliance/dlp-scanner) in the compliance package.
Choose the bucket first:

1. **Can a checksum or registry validate it?** Write the candidate regex,
   then require the validator (Luhn-style, mod-97-style, an issuer range
   table) before the match counts — this is the only shape allowed to raise
   a `block` verdict on a send gate.
2. **Is shape alone ambiguous?** Anchor it with a context-word dictionary
   gate, the same way the passport bucket requires a `passport` cue plus at
   least one digit in the token.
3. **Neither works?** Keep it out of the send gate. A category that can only
   be fuzzy-matched belongs in a storage-side scrubber, not in a rule a
   strict-mode organization will hard-block on.

Whichever bucket you pick, keep the two invariants every shipped bucket
holds: **findings carry category + offsets only** (a new author must never
echo the matched value into a violation message, header, or audit event),
and the scan stays **pure and synchronous** (no DB, no network — the linter
runs it per keystroke). A change to the shipped categories goes through
support today, as described on the feature page — treat new-bucket requests
the same way.

***

## 7. Reading it with the rest of the map

| View | Page | What it answers |
| - | - | - |
| Configuration & enforcement | [DLP scanner](/compliance/dlp-scanner) | Mode matrix, category floor, response shapes, audit |
| Pipeline plumbing | [The policy-scan pipeline model](/concepts/policy-scan-pipeline-model) | Hook placement, verdict modes, header contract, fail-closed boundary |
| Registry of every gate rule | [Pre-send policy scanner](/compliance/policy-scanner) | The full rule catalog the DLP buckets roll up into |
| Sibling scope | [Inbox AI privacy](/compliance/inbox-ai-privacy) | The LLM-boundary gates the DLP rule never depends on |
