Skip to main content

Tool-call trajectory grading

The trajectory grader runs over the cumulative tool-call list of an agent conversation and checks three things: argument correctness, call order, and forbidden-call violations. It is a deterministic evaluator, not an LLM judge, so it gives the same verdict every time the same calls are replayed. Use it when you want to gate a rollout on facts the agent controls directly — “did it look up the account before issuing a refund?” — rather than on style or open-ended quality. It complements the keyword, substring, and LLM-judge scoring covered in Simulation & Regression Eval Suite and Voice eval runs.

1. What trajectory grading checks

Every recorded tool call in a conversation has a tool name and a sanitized input payload. The grader evaluates that sequence against assertions you define in the scenario or dataset row:
  • Argument match (tool_call_args) — a named tool must have been called with arguments that match an exact, subset, or regex rule. You can pin the check to the Nth call of that tool.
  • Order (expected_tool_sequence) — an ordered list of tool steps must appear, in order, as a subsequence of the trace. Other calls may interleave without breaking the check.
  • Forbidden call (allowed_tools) — any tool call whose name is not in the allowlist fails the run.
These checks run in addition to the older presence-only checks (expected_tool_calls: “was this tool called at any point?”). A scenario can mix trajectory assertions, keyword assertions, and LLM-judge rubrics.

2. Where to use trajectory assertions

Two surfaces accept the same assertion shape:
  • Agent Studio Testing tab — dry-run, batch-simulation, voice-simulation, and persona-simulation scenarios all accept tool_call_args, expected_tool_sequence, and allowed_tools.
  • Agent evals datasets — add trajectory assertions to a dataset row so that an eval run grades the agent’s tool behaviour at run completion.
For the full eval-run loop, see Agent evals: datasets, runs, and pass-rate gates. For promoting a real production conversation into a dataset row, see Promote production traces into eval datasets.

3. Assertion types

All three assertion types are optional and additive. Omit one and it is a no-op.

Argument match

A tool_call_args entry asserts that a specific tool was called with a specific input shape.
mode may be one of: Add call_index (1-based) to pin the assertion to the Nth invocation of that tool. Without call_index, any matching invocation is enough.

Order

expected_tool_sequence asserts that a list of tools fired in a specific order. The check looks for an ordered subsequence, so unrelated calls can appear in between.
Each step can also carry a match object to require both order and arguments at that position.

Forbidden call

allowed_tools is a deny-by-default allowlist. Any observed tool whose name is not in the list fails the run.
This is useful for safety checks, such as confirming that a frontline support agent never calls a human-escalation or admin tool.

4. Running an eval

Trajectory assertions are evaluated when an eval run completes. The runner:
  1. Executes the conversation against the agent.
  2. Collects the full, ordered list of tool calls and their persisted inputs.
  3. Runs the three assertion families over that list.
  4. Appends the results to the scenario or run’s assertion list.
You do not call the grader directly; you declare the assertions in the scenario, dataset row, or simulation payload and read the results in the run detail.

5. Reading results

Each assertion produces one result with these fields:
  • name identifies the assertion. Argument-match results are named tool_call_args:<tool> or tool_call_args:<tool>#<call_index>. Sequence results are named tool_sequence. Forbidden-call results are named no_unexpected_tool.
  • status is pass or fail.
  • reason is empty on pass and describes the mismatch on fail.
Failures roll up to the run gate the same way other assertion failures do: a failed trajectory assertion lowers the pass rate and, depending on the scenario’s min_pass_rate, can block the run.

6. Interleaving caveats

The sequence check is intentionally lenient about interleaved calls. It treats the trace as an ordered subsequence problem, matching the way production agents often emit retrieval, logging, or state-check calls between the steps that matter to the test. For example, this trace passes the sequence [lookup_account, issue_refund] because the relative order is correct:
The grader does not care about adjacency, repetition, or the total number of calls — only that every required step appears, in order, somewhere in the trace. If you need the Nth occurrence of a specific tool to carry specific arguments, use tool_call_args with call_index instead of, or alongside, the sequence assertion.

7. Composing with persona simulation and CI gating

Trajectory assertions work inside the persona-simulation harness. Add them to a scenario alongside the LLM-judge rubric so that one run checks both behaviour and tool correctness:
  • The LLM judge scores the transcript for tone, policy adherence, and outcome.
  • The trajectory grader scores the exact tools and order.
A CI gate can fail if either layer reports a failure. Treat status: "fail" from any trajectory assertion as blocking, and read the reason field to diagnose whether the agent used the wrong arguments, the wrong order, or a forbidden tool.

8. Worked evals

Refund agent — argument correctness

A refund agent should issue a refund only after it has the right order and amount. The following assertion requires that issue_refund is called with order_id ord_42 and amount_cents 5000:
A call to issue_refund with the wrong amount fails with a reason such as “no call to issue_refund matched the expected arguments.”

Technical support — forbidden escalate

A tier-one technical support agent should not escalate to a human until it has run the standard diagnostic tools. This allowlist rejects any trace that calls something outside the approved set:
If the agent calls transfer_to_human before the diagnostics, the no_unexpected_tool result fails and names the forbidden tool.

Handoff agent — order constraint

A handoff agent must identify the customer, fetch the relevant context, and only then transfer. This sequence asserts that order:
If the agent transfers before fetching context, the tool_sequence result fails and reports how far the trace got through the expected order.

See also