Skip to main content

PII egress guard: the last-pass scrub on agent tool arguments

An AI agent’s PII controls are all output-side until this one: the model’s replies are scrubbed, tool results are scrubbed before re-entry, error surfaces are redacted. None of that touches the arguments of a tool call — the payload the model composes and the executor hands to an outbound connector. The PII egress guard closes that gap. It is an operator-controlled, opt-in flag on your agent’s safety configuration: your organization decides whether it runs, and the platform never flips the posture for you. This page is the model behind the flag: what the guard scrubs, the leak it closes, why the scrub is structure-preserving, which tools to enable it for, and why it can never turn a dispatch into an error.
The flag is a tenant-owned control. It lives in the agent’s safety configuration, defaults to off, and changes behavior only for the agents you enable it on.

1. What it is

When an agent run produces a tool call, the executor resolves the call’s arguments against the tool’s input schema and dispatches them to the connector behind that tool — a webhook endpoint, a CRM write, a messaging send, or a tool from your own MCP server. The egress guard is the last pass over those arguments before dispatch: with the flag on, every string value inside the argument payload is scanned for PII, and any hit is rewritten in place before the connector receives the payload. Two placement facts define the model:
  • It runs before dispatch, not after. The connector — the system that would carry the value out of your Orbit workspace — never sees the unscrubbed argument.
  • It is separate from output redaction. The output-side scrubber works on text the model emits for a user to read; the egress guard works on the structured argument object a tool consumes. Neither substitutes for the other.

2. The gap it closes

The existing guards form a complete-sounding perimeter: None of them touches a tool call’s input. The model composes those arguments from the conversation — and the conversation is typed by your end users. A user who says “email my records to me@example.com” or reads out an IBAN like DE44 5001 … gives the model a PII value it will faithfully place into the tool call. Without the egress guard, that value flows straight through the connector while every output-side guard reads clean — which is a real egress path under a data-minimization boundary like GDPR Art. 5(1)(c) or HIPAA’s minimum-necessary rule. The egress guard moves the scrub onto that last path: the same detector set that protects output now also protects the tool-call arguments, at the one point where they are still on your side of the connector boundary.

3. Structure-safe redaction

Tool arguments are structured JSON, not prose. A naive scrub — serializing the payload, redacting the string, and parsing it back — would rewrite keys, corrupt nested values, or throw on re-parse, and a corrupted argument object is worse than a leaked one. The guard instead walks the argument object recursively, string value by string value:
  • Objects and arrays are descended into; keys, numbers, booleans, and nulls are never touched.
  • Each string value is scanned independently; only values containing a PII hit are rewritten.
  • When nothing matches, the original object is returned by identity — so the serialized payload is byte-for-byte identical and downstream size checks, sandbox envelopes, and audit rows see no delta on the no-PII path.
Before and after, for a webhook-style tool call:
The shape survives; only the values that carried PII change.

4. Opt-in posture — and which tools tolerate it

The guard is opt-in: it runs only when the agent’s safety configuration sets pii_egress: true. Agents without the flag behave exactly as before.
Opt-in rather than always-on because argument structure is not generic-safe across every tool. Some tools parse a scrubbed string value into a stricter format, and a [REDACTED] sentinel breaks that parse — the canonical example is the orbit_send_sms tool, which expects its to field to parse as an E.164 phone number. Redacting an E.164 hit there turns a dispatchable call into a failed one. The practical rule:
  • Enable it for agents whose tools carry free-text payloads — webhook connectors, CRM field writes, and tools on your own MCP server. These treat arguments as opaque strings and tolerate a sentinel.
  • Leave it off for agents whose tools parse arguments into structured formats — phone recipients, typed amount fields, strict identifiers. Prefer the output-side guards and approval gates there.
The guard scans with the universal floor of PII detectors — email addresses, E.164 phone numbers, US social-security numbers, and credit-card numbers — with no locale bundle required. This is the same always-on detector set every existing tool-output guard uses, so enabling the flag never widens coverage beyond what your output side already treats as PII. If your agent has opted into locale-specific detector bundles on the output side, those locale codes pass through to the egress guard as well — one consistent pattern set across the inbound and outbound surfaces, rather than two diverging definitions of what counts as PII.

6. Fail-open posture

The guard is fail-open by design. If the detector throws while scanning a value, that value passes through to the connector unchanged and a warning is logged. A dispatch is never turned into an error by the scrub: the guard’s job is to remove PII when it can, and it never holds a tool call hostage to its own failure. The same posture applies to the no-PII path — with nothing matched, the payload the connector receives is the payload the model composed, without an extra serialize cycle in between.

Reading it with the rest of the map