PII egress guard: the last-pass scrub on agent tool arguments
An AI agent’s PII controls are all output-side until this one: the model’s replies are scrubbed, tool results are scrubbed before re-entry, error surfaces are redacted. None of that touches the arguments of a tool call — the payload the model composes and the executor hands to an outbound connector. The PII egress guard closes that gap. It is an operator-controlled, opt-in flag on your agent’s safety configuration: your organization decides whether it runs, and the platform never flips the posture for you. This page is the model behind the flag: what the guard scrubs, the leak it closes, why the scrub is structure-preserving, which tools to enable it for, and why it can never turn a dispatch into an error.The flag is a tenant-owned control. It lives in the agent’s safety
configuration, defaults to off, and changes behavior only for the agents
you enable it on.
1. What it is
When an agent run produces a tool call, the executor resolves the call’s arguments against the tool’s input schema and dispatches them to the connector behind that tool — a webhook endpoint, a CRM write, a messaging send, or a tool from your own MCP server. The egress guard is the last pass over those arguments before dispatch: with the flag on, every string value inside the argument payload is scanned for PII, and any hit is rewritten in place before the connector receives the payload. Two placement facts define the model:- It runs before dispatch, not after. The connector — the system that would carry the value out of your Orbit workspace — never sees the unscrubbed argument.
- It is separate from output redaction. The output-side scrubber works on text the model emits for a user to read; the egress guard works on the structured argument object a tool consumes. Neither substitutes for the other.
2. The gap it closes
The existing guards form a complete-sounding perimeter:
None of them touches a tool call’s input. The model composes those
arguments from the conversation — and the conversation is typed by your end
users. A user who says “email my records to me@example.com” or reads out
an IBAN like DE44 5001 … gives the model a PII value it will faithfully
place into the tool call. Without the egress guard, that value flows
straight through the connector while every output-side guard reads clean —
which is a real egress path under a data-minimization boundary like GDPR
Art. 5(1)(c) or HIPAA’s minimum-necessary rule.
The egress guard moves the scrub onto that last path: the same detector set
that protects output now also protects the tool-call arguments, at the one
point where they are still on your side of the connector boundary.
3. Structure-safe redaction
Tool arguments are structured JSON, not prose. A naive scrub — serializing the payload, redacting the string, and parsing it back — would rewrite keys, corrupt nested values, or throw on re-parse, and a corrupted argument object is worse than a leaked one. The guard instead walks the argument object recursively, string value by string value:- Objects and arrays are descended into; keys, numbers, booleans, and nulls are never touched.
- Each string value is scanned independently; only values containing a PII hit are rewritten.
- When nothing matches, the original object is returned by identity — so the serialized payload is byte-for-byte identical and downstream size checks, sandbox envelopes, and audit rows see no delta on the no-PII path.
4. Opt-in posture — and which tools tolerate it
The guard is opt-in: it runs only when the agent’s safety configuration setspii_egress: true. Agents without the flag behave exactly as before.
[REDACTED] sentinel breaks that parse — the canonical
example is the orbit_send_sms tool, which expects its to field to parse
as an E.164 phone number. Redacting an E.164 hit there turns a dispatchable
call into a failed one.
The practical rule:
- Enable it for agents whose tools carry free-text payloads — webhook connectors, CRM field writes, and tools on your own MCP server. These treat arguments as opaque strings and tolerate a sentinel.
- Leave it off for agents whose tools parse arguments into structured formats — phone recipients, typed amount fields, strict identifiers. Prefer the output-side guards and approval gates there.