Tool-call trajectory grading
The trajectory grader runs over the cumulative tool-call list of an agent conversation and checks three things: argument correctness, call order, and forbidden-call violations. It is a deterministic evaluator, not an LLM judge, so it gives the same verdict every time the same calls are replayed. Use it when you want to gate a rollout on facts the agent controls directly — “did it look up the account before issuing a refund?” — rather than on style or open-ended quality. It complements the keyword, substring, and LLM-judge scoring covered in Simulation & Regression Eval Suite and Voice eval runs.1. What trajectory grading checks
Every recorded tool call in a conversation has a tool name and a sanitized input payload. The grader evaluates that sequence against assertions you define in the scenario or dataset row:- Argument match (
tool_call_args) — a named tool must have been called with arguments that match anexact,subset, orregexrule. You can pin the check to the Nth call of that tool. - Order (
expected_tool_sequence) — an ordered list of tool steps must appear, in order, as a subsequence of the trace. Other calls may interleave without breaking the check. - Forbidden call (
allowed_tools) — any tool call whose name is not in the allowlist fails the run.
expected_tool_calls: “was this tool called at any point?”). A scenario can
mix trajectory assertions, keyword assertions, and LLM-judge rubrics.
2. Where to use trajectory assertions
Two surfaces accept the same assertion shape:- Agent Studio Testing tab — dry-run, batch-simulation, voice-simulation,
and persona-simulation scenarios all accept
tool_call_args,expected_tool_sequence, andallowed_tools. - Agent evals datasets — add trajectory assertions to a dataset row so that an eval run grades the agent’s tool behaviour at run completion.
3. Assertion types
All three assertion types are optional and additive. Omit one and it is a no-op.Argument match
Atool_call_args entry asserts that a specific tool was called with a
specific input shape.
mode may be one of:
Add
call_index (1-based) to pin the assertion to the Nth invocation of that
tool. Without call_index, any matching invocation is enough.
Order
expected_tool_sequence asserts that a list of tools fired in a specific order.
The check looks for an ordered subsequence, so unrelated calls can appear in
between.
match object to require both order and arguments
at that position.
Forbidden call
allowed_tools is a deny-by-default allowlist. Any observed tool whose name is
not in the list fails the run.
4. Running an eval
Trajectory assertions are evaluated when an eval run completes. The runner:- Executes the conversation against the agent.
- Collects the full, ordered list of tool calls and their persisted inputs.
- Runs the three assertion families over that list.
- Appends the results to the scenario or run’s assertion list.
5. Reading results
Each assertion produces one result with these fields:nameidentifies the assertion. Argument-match results are namedtool_call_args:<tool>ortool_call_args:<tool>#<call_index>. Sequence results are namedtool_sequence. Forbidden-call results are namedno_unexpected_tool.statusispassorfail.reasonis empty on pass and describes the mismatch on fail.
min_pass_rate, can block the run.
6. Interleaving caveats
The sequence check is intentionally lenient about interleaved calls. It treats the trace as an ordered subsequence problem, matching the way production agents often emit retrieval, logging, or state-check calls between the steps that matter to the test. For example, this trace passes the sequence[lookup_account, issue_refund]
because the relative order is correct:
tool_call_args with call_index instead of, or alongside,
the sequence assertion.
7. Composing with persona simulation and CI gating
Trajectory assertions work inside the persona-simulation harness. Add them to a scenario alongside the LLM-judge rubric so that one run checks both behaviour and tool correctness:- The LLM judge scores the transcript for tone, policy adherence, and outcome.
- The trajectory grader scores the exact tools and order.
status: "fail"
from any trajectory assertion as blocking, and read the reason field to
diagnose whether the agent used the wrong arguments, the wrong order, or a
forbidden tool.
8. Worked evals
Refund agent — argument correctness
A refund agent should issue a refund only after it has the right order and amount. The following assertion requires thatissue_refund is called with
order_id ord_42 and amount_cents 5000:
issue_refund with the wrong amount fails with a reason such as
“no call to issue_refund matched the expected arguments.”
Technical support — forbidden escalate
A tier-one technical support agent should not escalate to a human until it has run the standard diagnostic tools. This allowlist rejects any trace that calls something outside the approved set:transfer_to_human before the diagnostics, the
no_unexpected_tool result fails and names the forbidden tool.
Handoff agent — order constraint
A handoff agent must identify the customer, fetch the relevant context, and only then transfer. This sequence asserts that order:tool_sequence result
fails and reports how far the trace got through the expected order.
See also
- Simulation & Regression Eval Suite — scripted multi-turn scenarios with LLM-judge and keyword assertions.
- Voice eval runs — reading voice-eval run results.
- Continuous production evals — production-sampled eval runs.
- Promote production traces into eval datasets — turn a real conversation into a dataset row.