Skip to main content

Guardrail policies and red-team gates: how agents are governed

An agent you can reason about ends at promotion. A prompt change, a model swap, or a new golden example set is never just “applied” — it is versioned, scored against a pinned correctness suite, and optionally held behind a human approval and a pre-deploy red-team gate. The on-demand red-team run and the guardrail-analytics rollup answer “is the live agent safety-kept today?”; the promotion pipeline answers “did this candidate make the agent worse than the pinned baseline?” This page collects the four surfaces that together make that question answerable: the guardrail policy library, the token-budget resolver, the red-team run lifecycle, and the promote-time gates. Trust anchors — the human-approval and the AI-disclosure ledger — close the page.

1. Guardrail policy authoring and token budgets

Policies: version once, apply to many agents

Every agent carries its own inline guardrails blob (config.guardrails plus safety_config), tuned per agent. When several customer-facing agents share one safety posture — PII / credential blocking, topic and word denylists, KB-citation enforcement, a factual-consistency threshold — you can model that posture as a guardrail policy: a named, versioned object from the built-in policy library, then apply it to each agent that should inherit it. The public surface is GET /api/v1/agents/guardrail-policies (the library, grouped by tier) and GET /api/v1/agents/guardrail-policies/:policyId (one policy including the full safety_config the apply would merge, so the dashboard previews the exact change). Applying uses POST /api/v1/agents/:id/guardrail-policy, and the provenance read-back is GET /api/v1/agents/:id/guardrail-policy — it reports the stamped { id, name, version, applied_at } and whether that version is stale relative to the current library, so the dashboard can prompt a re-apply. Applying a policy merges its safety_config patch onto the agent’s existing safety_config — keys the policy does not govern survive — and stamps config.guardrails.applied_policy so the agent records which policy and version it was configured from. Drift is visible rather than silent. The policy is a snapshot, not a live reference: new library versions only engage when you re-apply, matching how platform-versioned templates behave.

Token budgets and the downgrade ladder

Guardrails attach to cost as well as content. Three cent-denominated limits already exist — a per-conversation cost cap, an org-wide daily/monthly budget downshift, and a per-turn max_tokens ceiling. What token_budget and downgrade_ladder add is a per-agent, token-denominated posture with progressive enforcement instead of a hard stop:
  • token_budget — a per-conversation token allowance.
  • downgrade_ladder.steps — an ordered map of utilization thresholds to cheaper models, so an agent steps down as it approaches its downgrade_ladder.daily_token_cap rather than staying on its primary model until a hard stop.
The resolved decision — “which model should the next turn use, and how much of the budget or cap is consumed?” — is fetched in one round trip by the agent-runtime on the hot pre-turn path via the internal GET /api/v1/agents/:id/token-guardrail route. The resolver never throws. An unset, malformed, or unusable ladder simply means “no downgrade configured” — a breached cap is reported, and the agent keeps its primary model rather than ever resolving to an unvalidated model id. Bad shapes and unsupported model ids are rejected with a 422 at save time.

2. The red-team run lifecycle and the pre-deploy gate

On-demand runs: replay the adversarial pack in sandbox mode

Monitoring a live agent for drifting safety is only half the story — the other half is an opt-in red-team gate on the promote path. Both halves replay the same built-in adversarial pack: twelve probes spanning jailbreak, prompt injection, data exfiltration, and policy-guardrail probes. The on-demand surface is GET /api/v1/agents/:id/red-team/pack (the probe catalogue with id, category, title, and expectation; raw compromise markers are intentionally omitted so the canary tokens used for detection are never advertised in an API response) and POST /api/v1/agents/:id/red-team/run (replay the full pack, or a category subset, against the live agent in sandbox mode and return a safety scorecard). Each probe replays as one sandbox chat call to the agent-runtime: the probe’s setup turns become replay history and the attack is the current message. Sandbox short-circuits mutative tools, skips billing, and does not write to memory — so a red-team run cannot alter tenant-side state. A transport or non-2xx failure is recorded as an error, not a compromise — a runtime blip is not evidence the agent leaked. The scorecard summarises resisted, compromised, and errored probes into an overall safety score and grade.

The gate: pinned baseline plus floor on the promote path

The gate closes the loop the on-demand scorecard cannot: blocking an unsafe version before it reaches production. Where the on-demand scorer answers “is the live agent safe today?”, the gate answers “does this candidate version regress the agent’s pinned safety baseline?” Gate config and the pinned baseline live on organization settings — the same non-versioned location as the promotion regression gate and the prompt- promotion-approval policy. An agent is promoted only when its gate is explicitly enabled; default is off, and an agent with no gate promotes exactly as it did before. The gate replays the full pack against the candidate version — the version under test, not the live one — by driving the LLM leg directly under that version’s frozen system prompt, model, and sampling. It blocks when either of two conditions holds:
  • the candidate’s overall safety score falls below the configured min_safety_score floor — which bites on the very first run, before any baseline exists — or
  • at least one probe that the pinned baseline had resisted now reads compromised (a genuine new regression; a probe the baseline already had compromised is not double-counted).
On a pass the new scorecard becomes the next pinned baseline, so gradual drift is caught incrementally rather than only against the very first run. The gate is fail-closed: a gate that crashes blocks the promote rather than waving it through, because a safety gate that fails open is worse than none. The promote response carries the gate’s report — blocked reasons, scorecard, and the pinned baseline it was compared against — as audit evidence. The promote path composes the same discipline as every other pre-deploy gate: an org-level separation-of-duties check, the prompt-promotion approval policy when configured, then the regression gate, then the red-team gate — each evaluated outside the promote transaction, because the gate’s LLM calls spend real pre-production tokens and the transacted flip must be atomic.

3. Guardrail analytics: violations by category and agent

Configuring guardrails without being able to see how they behave on live traffic is half a control. Every completed agent turn persists a governance row in the tenant’s AI decision log (ai_turn_audit), carrying the outcome — refused when the guardrails or compliance stopped the turn — plus the factual-grounding confidence. The tenant-wide rollup GET /api/v1/agents/guardrail-analytics aggregates those rows so an operator can see, per agent, which violation categories fire most and which agents are affected, plus grounding stats and a daily trend. Optional filters narrow by agentId, from / to, a days window, a lowConfidenceThreshold, and agentLimit. The response returns outcome totals and firing rates, a per-agent breakdown each carrying its applied policy, a per-policy breakdown (“which policies fire most”), grounding stats, and a daily trend. The aggregate surfaces the whole tenant’s AI decision behaviour, so the route is read-only, requires owner or admin, and sits behind the matching tenant-wide export gate. Operators tune guardrails by watching which category fires and per which agent; teams validate that a policy version actually lowers refusal of the categories it governs before they widen its application.

4. The promotion pipeline: gates compose into one verdict

Promotion to a live agent is the point where guardrails, eval, and human review meet. The same pre-deploy red-team gate that blocks unsafe prompt or model changes composes with the correctness regression gate and the prompt- promotion approval policy over the exact same promote endpoint — POST /agents/:id/versions/:vid/promote and its prompt-rollback alias. Correctness regression gates compare the candidate’s pinned golden-set suite against the last passing baseline. The red-team gate compares the candidate’s safety scorecard against the pinned safety baseline. Experiments add an auto-eval rollup on the variants: because a variant’s eval score is one join away from its already-versioned prompt id, the completed agent_eval_runs rollup yields one informational quality signal the results panel can show alongside the conversion and cost signals. That auto-eval signal is deliberately informational-only — it never gates the winner automatically, so a human judge always decides promotion over a composite view of lift, cost, and quality. Together these let a builder reason about governance budgets before exporting a prompt set: the guardrails a bundle carries, the cost envelope it is held under, the safety verdict it must preserve, and the mechanism that blocks a regression before customers see it.

5. Trust anchors

Governance depends on human and evidentiary rails. The action approval trails the call that turns a tool into a proposal — pending action, exact arguments, estimated cost — so a person, not a model, decides privileged tool calls. The AI-disclosure ledger folds the agent’s turn-by-turn decision log and the tenant’s disclosure posture into a signed provenance record that answers when and where an AI — not a human — spoke, and proves the record was not edited. Where governance gates block an unsafe change, the approval gate blocks a privileged tool call, and the disclosure ledger binds the visible effect of whichever gate allowed the turn through.