> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orbit.devotel.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Tool-call trajectory grading

> Assert which tools an agent called, in what order, and with which arguments — deterministic eval grader for agent tool behaviour.

# Tool-call trajectory grading

The trajectory grader runs over the cumulative tool-call list of an agent
conversation and checks three things: **argument correctness**, **call order**,
and **forbidden-call violations**. It is a deterministic evaluator, not an LLM
judge, so it gives the same verdict every time the same calls are replayed.

Use it when you want to gate a rollout on facts the agent controls directly —
"did it look up the account before issuing a refund?" — rather than on style or
open-ended quality. It complements the keyword, substring, and LLM-judge scoring
covered in [Simulation & Regression Eval Suite](/agents/simulation-eval-suite)
and [Voice eval runs](/agents/voice-eval-runs).

## 1. What trajectory grading checks

Every recorded tool call in a conversation has a tool name and a sanitized
input payload. The grader evaluates that sequence against assertions you define
in the scenario or dataset row:

* **Argument match** (`tool_call_args`) — a named tool must have been called
  with arguments that match an `exact`, `subset`, or `regex` rule. You can pin
  the check to the Nth call of that tool.
* **Order** (`expected_tool_sequence`) — an ordered list of tool steps must
  appear, in order, as a subsequence of the trace. Other calls may interleave
  without breaking the check.
* **Forbidden call** (`allowed_tools`) — any tool call whose name is not in the
  allowlist fails the run.

These checks run in addition to the older presence-only checks
(`expected_tool_calls`: "was this tool called at any point?"). A scenario can
mix trajectory assertions, keyword assertions, and LLM-judge rubrics.

## 2. Where to use trajectory assertions

Two surfaces accept the same assertion shape:

* **Agent Studio Testing tab** — dry-run, batch-simulation, voice-simulation,
  and persona-simulation scenarios all accept `tool_call_args`,
  `expected_tool_sequence`, and `allowed_tools`.
* **Agent evals datasets** — add trajectory assertions to a dataset row so that
  an eval run grades the agent's tool behaviour at run completion.

For the full eval-run loop, see
[Agent evals: datasets, runs, and pass-rate gates](/guides/agent-evals-datasets-loop).
For promoting a real production conversation into a dataset row, see
[Promote production traces into eval datasets](/guides/agent-trace-to-dataset).

## 3. Assertion types

All three assertion types are optional and additive. Omit one and it is a
no-op.

### Argument match

A `tool_call_args` entry asserts that a specific tool was called with a
specific input shape.

```json theme={null}
{
  "tool_call_args": [
    {
      "tool": "issue_refund",
      "match": {
        "mode": "subset",
        "args": { "order_id": "ord_42", "amount_cents": 5000 }
      }
    }
  ]
}
```

`mode` may be one of:

| Mode | Meaning |
| - | - |
| `exact` | The persisted tool input must deep-equal `args`. Extra keys fail the check. |
| `subset` | Every key in `args` must deep-equal the same key on the persisted input. Extra keys on the actual input are ignored. |
| `regex` | The regular expression in `pattern` is tested against `JSON.stringify(input)`. Optional `flags` are supported. |

Add `call_index` (1-based) to pin the assertion to the Nth invocation of that
tool. Without `call_index`, any matching invocation is enough.

```json theme={null}
{
  "tool_call_args": [
    {
      "tool": "kb_query",
      "call_index": 2,
      "match": {
        "mode": "regex",
        "pattern": "refund policy",
        "flags": "i"
      }
    }
  ]
}
```

### Order

`expected_tool_sequence` asserts that a list of tools fired in a specific order.
The check looks for an ordered subsequence, so unrelated calls can appear in
between.

```json theme={null}
{
  "expected_tool_sequence": [
    { "tool": "lookup_account" },
    { "tool": "check_refund_eligibility" },
    { "tool": "issue_refund" }
  ]
}
```

Each step can also carry a `match` object to require both order and arguments
at that position.

```json theme={null}
{
  "expected_tool_sequence": [
    { "tool": "lookup_account" },
    {
      "tool": "issue_refund",
      "match": { "mode": "subset", "args": { "amount_cents": 5000 } }
    }
  ]
}
```

### Forbidden call

`allowed_tools` is a deny-by-default allowlist. Any observed tool whose name is
not in the list fails the run.

```json theme={null}
{
  "allowed_tools": ["lookup_account", "check_refund_eligibility", "issue_refund"]
}
```

This is useful for safety checks, such as confirming that a frontline support
agent never calls a human-escalation or admin tool.

## 4. Running an eval

Trajectory assertions are evaluated when an eval run completes. The runner:

1. Executes the conversation against the agent.
2. Collects the full, ordered list of tool calls and their persisted inputs.
3. Runs the three assertion families over that list.
4. Appends the results to the scenario or run's assertion list.

You do not call the grader directly; you declare the assertions in the
scenario, dataset row, or simulation payload and read the results in the run
detail.

## 5. Reading results

Each assertion produces one result with these fields:

```json theme={null}
{
  "name": "tool_call_args:issue_refund",
  "status": "pass",
  "reason": ""
}
```

* `name` identifies the assertion. Argument-match results are named
  `tool_call_args:<tool>` or `tool_call_args:<tool>#<call_index>`. Sequence
  results are named `tool_sequence`. Forbidden-call results are named
  `no_unexpected_tool`.
* `status` is `pass` or `fail`.
* `reason` is empty on pass and describes the mismatch on fail.

Failures roll up to the run gate the same way other assertion failures do: a
failed trajectory assertion lowers the pass rate and, depending on the
scenario's `min_pass_rate`, can block the run.

## 6. Interleaving caveats

The sequence check is intentionally lenient about interleaved calls. It treats
the trace as an ordered subsequence problem, matching the way production
agents often emit retrieval, logging, or state-check calls between the steps
that matter to the test.

For example, this trace passes the sequence `[lookup_account, issue_refund]`
because the relative order is correct:

```json theme={null}
[
  { "tool": "lookup_account", "input": { "customer_id": "cus_1" } },
  { "tool": "log_event", "input": { "event": "account_viewed" } },
  { "tool": "kb_query", "input": { "q": "refund policy" } },
  { "tool": "issue_refund", "input": { "order_id": "ord_42" } }
]
```

The grader does not care about adjacency, repetition, or the total number of
calls — only that every required step appears, in order, somewhere in the
trace. If you need the Nth occurrence of a specific tool to carry specific
arguments, use `tool_call_args` with `call_index` instead of, or alongside,
the sequence assertion.

## 7. Composing with persona simulation and CI gating

Trajectory assertions work inside the
[persona-simulation harness](/agents/persona-simulation). Add them to a
scenario alongside the LLM-judge rubric so that one run checks both behaviour
and tool correctness:

* The **LLM judge** scores the transcript for tone, policy adherence, and
  outcome.
* The **trajectory grader** scores the exact tools and order.

A CI gate can fail if either layer reports a failure. Treat `status: "fail"`
from any trajectory assertion as blocking, and read the `reason` field to
diagnose whether the agent used the wrong arguments, the wrong order, or a
forbidden tool.

## 8. Worked evals

### Refund agent — argument correctness

A refund agent should issue a refund only after it has the right order and
amount. The following assertion requires that `issue_refund` is called with
`order_id` `ord_42` and `amount_cents` `5000`:

```json theme={null}
{
  "tool_call_args": [
    {
      "tool": "issue_refund",
      "match": {
        "mode": "subset",
        "args": { "order_id": "ord_42", "amount_cents": 5000 }
      }
    }
  ]
}
```

A call to `issue_refund` with the wrong amount fails with a reason such as
"no call to `issue_refund` matched the expected arguments."

### Technical support — forbidden escalate

A tier-one technical support agent should not escalate to a human until it has
run the standard diagnostic tools. This allowlist rejects any trace that calls
something outside the approved set:

```json theme={null}
{
  "allowed_tools": [
    "lookup_account",
    "run_diagnostics",
    "create_ticket",
    "kb_query"
  ]
}
```

If the agent calls `transfer_to_human` before the diagnostics, the
`no_unexpected_tool` result fails and names the forbidden tool.

### Handoff agent — order constraint

A handoff agent must identify the customer, fetch the relevant context, and
only then transfer. This sequence asserts that order:

```json theme={null}
{
  "expected_tool_sequence": [
    { "tool": "identify_customer" },
    { "tool": "fetch_case_context" },
    { "tool": "transfer_to_agent" }
  ]
}
```

If the agent transfers before fetching context, the `tool_sequence` result
fails and reports how far the trace got through the expected order.

## See also

* [Simulation & Regression Eval Suite](/agents/simulation-eval-suite) —
  scripted multi-turn scenarios with LLM-judge and keyword assertions.
* [Voice eval runs](/agents/voice-eval-runs) — reading voice-eval run results.
* [Continuous production evals](/agents/continuous-production-evals) —
  production-sampled eval runs.
* [Promote production traces into eval datasets](/guides/agent-trace-to-dataset)
  — turn a real conversation into a dataset row.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.