> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orbit.devotel.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Guardrail policies and red-team gates: how agents are governed

> The governance spine behind versioned AI agents — a reusable, versioned guardrail policy library, token-budget enforcement, the adversarial red-team pack, and the promote-time gates that reuse red-team verdicts to block an unsafe version before it goes live.

# Guardrail policies and red-team gates: how agents are governed

An agent you can reason about ends at promotion. A prompt change, a model
swap, or a new golden example set is never just "applied" — it is versioned,
scored against a pinned correctness suite, and optionally held behind a human
approval and a pre-deploy red-team gate. The on-demand red-team run and the
guardrail-analytics rollup answer "is the live agent safety-kept today?"; the
promotion pipeline answers "did this candidate make the agent worse than the
pinned baseline?"

This page collects the four surfaces that together make that question answerable:
the guardrail policy library, the token-budget resolver, the red-team run
lifecycle, and the promote-time gates. Trust anchors — the human-approval and
the AI-disclosure ledger — close the page.

## 1. Guardrail policy authoring and token budgets

### Policies: version once, apply to many agents

Every agent carries its own inline guardrails blob (`config.guardrails` plus
`safety_config`), tuned per agent. When several customer-facing agents share
one safety posture — PII / credential blocking, topic and word denylists,
KB-citation enforcement, a factual-consistency threshold — you can model that
posture as a **guardrail policy**: a named, versioned object from the built-in
policy library, then apply it to each agent that should inherit it.

The public surface is `GET /api/v1/agents/guardrail-policies` (the library,
grouped by tier) and `GET /api/v1/agents/guardrail-policies/:policyId` (one
policy including the full `safety_config` the apply would merge, so the
dashboard previews the exact change). Applying uses
`POST /api/v1/agents/:id/guardrail-policy`, and the provenance read-back is
`GET /api/v1/agents/:id/guardrail-policy` — it reports the stamped
`{ id, name, version, applied_at }` and whether that version is stale
relative to the current library, so the dashboard can prompt a re-apply.

Applying a policy merges its `safety_config` patch onto the agent's existing
`safety_config` — keys the policy does not govern survive — and stamps
`config.guardrails.applied_policy` so the agent records which policy and
version it was configured from. Drift is visible rather than silent. The
policy is a snapshot, not a live reference: new library versions only engage
when you re-apply, matching how platform-versioned templates behave.

### Token budgets and the downgrade ladder

Guardrails attach to cost as well as content. Three cent-denominated limits
already exist — a per-conversation cost cap, an org-wide daily/monthly budget
downshift, and a per-turn `max_tokens` ceiling. What `token_budget` and
`downgrade_ladder` add is a **per-agent, token-denominated** posture with
progressive enforcement instead of a hard stop:

* `token_budget` — a per-conversation token allowance.
* `downgrade_ladder.steps` — an ordered map of utilization thresholds to
  cheaper models, so an agent steps down as it approaches its
  `downgrade_ladder.daily_token_cap` rather than staying on its primary model
  until a hard stop.

The resolved decision — "which model should the next turn use, and how much of
the budget or cap is consumed?" — is fetched in one round trip by the
agent-runtime on the hot pre-turn path via the internal
`GET /api/v1/agents/:id/token-guardrail` route. The resolver never throws. An
unset, malformed, or unusable ladder simply means "no downgrade configured" —
a breached cap is reported, and the agent keeps its primary model rather than
ever resolving to an unvalidated model id. Bad shapes and unsupported model
ids are rejected with a 422 at save time.

## 2. The red-team run lifecycle and the pre-deploy gate

### On-demand runs: replay the adversarial pack in sandbox mode

Monitoring a live agent for drifting safety is only half the story — the
other half is an opt-in **red-team gate** on the promote path. Both halves
replay the same built-in adversarial pack: twelve probes spanning jailbreak,
prompt injection, data exfiltration, and policy-guardrail probes.

The on-demand surface is `GET /api/v1/agents/:id/red-team/pack` (the probe
catalogue with id, category, title, and expectation; raw compromise markers
are intentionally omitted so the canary tokens used for detection are never
advertised in an API response) and
`POST /api/v1/agents/:id/red-team/run` (replay the full pack, or a category
subset, against the live agent in sandbox mode and return a safety
scorecard).

Each probe replays as one sandbox chat call to the agent-runtime: the probe's
setup turns become replay history and the attack is the current message.
Sandbox short-circuits mutative tools, skips billing, and does not write to
memory — so a red-team run cannot alter tenant-side state. A transport or
non-2xx failure is recorded as an **error**, not a compromise — a runtime blip
is not evidence the agent leaked. The scorecard summarises resisted,
compromised, and errored probes into an overall safety score and grade.

### The gate: pinned baseline plus floor on the promote path

The gate closes the loop the on-demand scorecard cannot: blocking an unsafe
version **before** it reaches production. Where the on-demand scorer answers
"is the live agent safe today?", the gate answers "does this candidate version
regress the agent's pinned safety baseline?"

Gate config and the pinned baseline live on organization settings — the same
non-versioned location as the promotion regression gate and the prompt-
promotion-approval policy. An agent is promoted only when its gate is
explicitly enabled; default is off, and an agent with no gate promotes
exactly as it did before.

The gate replays the full pack against the candidate version — the version
under test, not the live one — by driving the LLM leg directly under that
version's frozen system prompt, model, and sampling. It blocks when either of
two conditions holds:

* the candidate's overall safety score falls below the configured
  `min_safety_score` floor — which bites on the very first run, before any
  baseline exists — or
* at least one probe that the pinned baseline had **resisted** now reads
  **compromised** (a genuine new regression; a probe the baseline already
  had compromised is not double-counted).

On a pass the new scorecard becomes the next pinned baseline, so gradual
drift is caught incrementally rather than only against the very first run.
The gate is fail-closed: a gate that crashes blocks the promote rather than
waving it through, because a safety gate that fails open is worse than none.
The promote response carries the gate's report — blocked reasons, scorecard,
and the pinned baseline it was compared against — as audit evidence.

The promote path composes the same discipline as every other pre-deploy
gate: an org-level separation-of-duties check, the prompt-promotion approval
policy when configured, then the regression gate, then the red-team gate —
each evaluated outside the promote transaction, because the gate's LLM calls
spend real pre-production tokens and the transacted flip must be atomic.

## 3. Guardrail analytics: violations by category and agent

Configuring guardrails without being able to see how they behave on live
traffic is half a control. Every completed agent turn persists a governance
row in the tenant's AI decision log (`ai_turn_audit`), carrying the outcome —
`refused` when the guardrails or compliance stopped the turn — plus the
factual-grounding confidence.

The tenant-wide rollup `GET /api/v1/agents/guardrail-analytics` aggregates
those rows so an operator can see, per agent, **which violation categories
fire most and which agents are affected**, plus grounding stats and a daily
trend. Optional filters narrow by `agentId`, `from` / `to`, a `days` window,
a `lowConfidenceThreshold`, and `agentLimit`. The response returns outcome
totals and firing rates, a per-agent breakdown each carrying its applied
policy, a per-policy breakdown ("which policies fire most"), grounding stats,
and a daily trend. The aggregate surfaces the whole tenant's AI decision
behaviour, so the route is read-only, requires `owner` or `admin`, and sits
behind the matching tenant-wide export gate.

Operators tune guardrails by watching which category fires and per which
agent; teams validate that a policy version actually lowers refusal of the
categories it governs before they widen its application.

## 4. The promotion pipeline: gates compose into one verdict

Promotion to a live agent is the point where guardrails, eval, and human
review meet. The same pre-deploy red-team gate that blocks unsafe prompt or
model changes composes with the correctness regression gate and the prompt-
promotion approval policy over the exact same promote endpoint —
`POST /agents/:id/versions/:vid/promote` and its `prompt-rollback` alias.

Correctness regression gates compare the candidate's pinned golden-set suite
against the last passing baseline. The red-team gate compares the candidate's
safety scorecard against the pinned safety baseline. Experiments add an
**auto-eval** rollup on the variants: because a variant's eval score is one
join away from its already-versioned prompt id, the completed
`agent_eval_runs` rollup yields one informational quality signal the results
panel can show alongside the conversion and cost signals. That auto-eval
signal is deliberately informational-only — it never gates the winner
automatically, so a human judge always decides promotion over a composite
view of lift, cost, and quality.

Together these let a builder reason about governance budgets before exporting
a prompt set: the guardrails a bundle carries, the cost envelope it is
held under, the safety verdict it must preserve, and the mechanism that
blocks a regression before customers see it.

## 5. Trust anchors

Governance depends on human and evidentiary rails. The **action approval**
trails the call that turns a tool into a proposal — pending action, exact
arguments, estimated cost — so a person, not a model, decides privileged
tool calls. The **AI-disclosure ledger** folds the agent's turn-by-turn
decision log and the tenant's disclosure posture into a signed provenance
record that answers when and where an AI — not a human — spoke, and proves
the record was not edited. Where governance gates block an unsafe change,
the approval gate blocks a privileged tool call, and the disclosure ledger
binds the visible effect of whichever gate allowed the turn through.

* [Action approvals: the human-in-the-loop gate](/concepts/action-approvals-model)
* [AI-disclosure ledger](/concepts/ai-disclosure-ledger)
* [Pick a testing harness: the decision table](/concepts/agent-testing-harness-picker)
* [AI agent architecture](/concepts/ai-agent-architecture)
