Guardrail policies and red-team gates: how agents are governed
An agent you can reason about ends at promotion. A prompt change, a model swap, or a new golden example set is never just “applied” — it is versioned, scored against a pinned correctness suite, and optionally held behind a human approval and a pre-deploy red-team gate. The on-demand red-team run and the guardrail-analytics rollup answer “is the live agent safety-kept today?”; the promotion pipeline answers “did this candidate make the agent worse than the pinned baseline?” This page collects the four surfaces that together make that question answerable: the guardrail policy library, the token-budget resolver, the red-team run lifecycle, and the promote-time gates. Trust anchors — the human-approval and the AI-disclosure ledger — close the page.1. Guardrail policy authoring and token budgets
Policies: version once, apply to many agents
Every agent carries its own inline guardrails blob (config.guardrails plus
safety_config), tuned per agent. When several customer-facing agents share
one safety posture — PII / credential blocking, topic and word denylists,
KB-citation enforcement, a factual-consistency threshold — you can model that
posture as a guardrail policy: a named, versioned object from the built-in
policy library, then apply it to each agent that should inherit it.
The public surface is GET /api/v1/agents/guardrail-policies (the library,
grouped by tier) and GET /api/v1/agents/guardrail-policies/:policyId (one
policy including the full safety_config the apply would merge, so the
dashboard previews the exact change). Applying uses
POST /api/v1/agents/:id/guardrail-policy, and the provenance read-back is
GET /api/v1/agents/:id/guardrail-policy — it reports the stamped
{ id, name, version, applied_at } and whether that version is stale
relative to the current library, so the dashboard can prompt a re-apply.
Applying a policy merges its safety_config patch onto the agent’s existing
safety_config — keys the policy does not govern survive — and stamps
config.guardrails.applied_policy so the agent records which policy and
version it was configured from. Drift is visible rather than silent. The
policy is a snapshot, not a live reference: new library versions only engage
when you re-apply, matching how platform-versioned templates behave.
Token budgets and the downgrade ladder
Guardrails attach to cost as well as content. Three cent-denominated limits already exist — a per-conversation cost cap, an org-wide daily/monthly budget downshift, and a per-turnmax_tokens ceiling. What token_budget and
downgrade_ladder add is a per-agent, token-denominated posture with
progressive enforcement instead of a hard stop:
token_budget— a per-conversation token allowance.downgrade_ladder.steps— an ordered map of utilization thresholds to cheaper models, so an agent steps down as it approaches itsdowngrade_ladder.daily_token_caprather than staying on its primary model until a hard stop.
GET /api/v1/agents/:id/token-guardrail route. The resolver never throws. An
unset, malformed, or unusable ladder simply means “no downgrade configured” —
a breached cap is reported, and the agent keeps its primary model rather than
ever resolving to an unvalidated model id. Bad shapes and unsupported model
ids are rejected with a 422 at save time.
2. The red-team run lifecycle and the pre-deploy gate
On-demand runs: replay the adversarial pack in sandbox mode
Monitoring a live agent for drifting safety is only half the story — the other half is an opt-in red-team gate on the promote path. Both halves replay the same built-in adversarial pack: twelve probes spanning jailbreak, prompt injection, data exfiltration, and policy-guardrail probes. The on-demand surface isGET /api/v1/agents/:id/red-team/pack (the probe
catalogue with id, category, title, and expectation; raw compromise markers
are intentionally omitted so the canary tokens used for detection are never
advertised in an API response) and
POST /api/v1/agents/:id/red-team/run (replay the full pack, or a category
subset, against the live agent in sandbox mode and return a safety
scorecard).
Each probe replays as one sandbox chat call to the agent-runtime: the probe’s
setup turns become replay history and the attack is the current message.
Sandbox short-circuits mutative tools, skips billing, and does not write to
memory — so a red-team run cannot alter tenant-side state. A transport or
non-2xx failure is recorded as an error, not a compromise — a runtime blip
is not evidence the agent leaked. The scorecard summarises resisted,
compromised, and errored probes into an overall safety score and grade.
The gate: pinned baseline plus floor on the promote path
The gate closes the loop the on-demand scorecard cannot: blocking an unsafe version before it reaches production. Where the on-demand scorer answers “is the live agent safe today?”, the gate answers “does this candidate version regress the agent’s pinned safety baseline?” Gate config and the pinned baseline live on organization settings — the same non-versioned location as the promotion regression gate and the prompt- promotion-approval policy. An agent is promoted only when its gate is explicitly enabled; default is off, and an agent with no gate promotes exactly as it did before. The gate replays the full pack against the candidate version — the version under test, not the live one — by driving the LLM leg directly under that version’s frozen system prompt, model, and sampling. It blocks when either of two conditions holds:- the candidate’s overall safety score falls below the configured
min_safety_scorefloor — which bites on the very first run, before any baseline exists — or - at least one probe that the pinned baseline had resisted now reads compromised (a genuine new regression; a probe the baseline already had compromised is not double-counted).
3. Guardrail analytics: violations by category and agent
Configuring guardrails without being able to see how they behave on live traffic is half a control. Every completed agent turn persists a governance row in the tenant’s AI decision log (ai_turn_audit), carrying the outcome —
refused when the guardrails or compliance stopped the turn — plus the
factual-grounding confidence.
The tenant-wide rollup GET /api/v1/agents/guardrail-analytics aggregates
those rows so an operator can see, per agent, which violation categories
fire most and which agents are affected, plus grounding stats and a daily
trend. Optional filters narrow by agentId, from / to, a days window,
a lowConfidenceThreshold, and agentLimit. The response returns outcome
totals and firing rates, a per-agent breakdown each carrying its applied
policy, a per-policy breakdown (“which policies fire most”), grounding stats,
and a daily trend. The aggregate surfaces the whole tenant’s AI decision
behaviour, so the route is read-only, requires owner or admin, and sits
behind the matching tenant-wide export gate.
Operators tune guardrails by watching which category fires and per which
agent; teams validate that a policy version actually lowers refusal of the
categories it governs before they widen its application.
4. The promotion pipeline: gates compose into one verdict
Promotion to a live agent is the point where guardrails, eval, and human review meet. The same pre-deploy red-team gate that blocks unsafe prompt or model changes composes with the correctness regression gate and the prompt- promotion approval policy over the exact same promote endpoint —POST /agents/:id/versions/:vid/promote and its prompt-rollback alias.
Correctness regression gates compare the candidate’s pinned golden-set suite
against the last passing baseline. The red-team gate compares the candidate’s
safety scorecard against the pinned safety baseline. Experiments add an
auto-eval rollup on the variants: because a variant’s eval score is one
join away from its already-versioned prompt id, the completed
agent_eval_runs rollup yields one informational quality signal the results
panel can show alongside the conversion and cost signals. That auto-eval
signal is deliberately informational-only — it never gates the winner
automatically, so a human judge always decides promotion over a composite
view of lift, cost, and quality.
Together these let a builder reason about governance budgets before exporting
a prompt set: the guardrails a bundle carries, the cost envelope it is
held under, the safety verdict it must preserve, and the mechanism that
blocks a regression before customers see it.