Skip to content

Evaluating and Monitoring LLM Applications · Evaluating RAG pipelines and agents · lesson 8 of 16 · 15 min

Evaluating agents: outcomes, trajectories and efficiency

Agents are harder to evaluate

An agent takes multiple steps: it plans, calls tools, reads results, and decides what to do next. Two runs on the same task can take different paths and both succeed, or take the same path and diverge at step six. You need to evaluate outcomes, trajectories and efficiency, usually in a controlled environment.

Three levels of agent evaluation

1. Outcome (end state). Did the task get done? Check the final state of the world, not the agent's claim: was the refund actually created with the right amount, did the file change, does the test pass, is the database row correct? Outcome checks are usually code-based and are the most important.

2. Trajectory (process). Was the path acceptable?

  • Were required steps taken (verify identity before revealing order details)?
  • Were forbidden actions avoided (no refund over the limit without approval)?
  • Were tool calls well-formed with correct arguments?
  • Did it recover from tool errors sensibly?
  • Did it ask for clarification when needed rather than guessing?

3. Efficiency. Steps, tool calls, tokens, wall-clock time and cost per successful task. Two agents with the same success rate can differ several-fold in cost.

Controlled environments

Agents act on the world, so evaluations need sandboxes: mock or containerized versions of tools and data that reset between runs. Options:

  • Mocked tools returning fixtures (fast, deterministic; good for unit-level agent tests).
  • Stateful simulators of your APIs and databases (for example a test instance seeded per run).
  • Containerized environments for coding or computer-use agents. The UK AI Security Institute's open-source Inspect framework, for example, supports tool-using agent evaluations with Docker-based sandboxes.
  • Simulated users: an LLM playing the customer, following a scenario script, for multi-turn tasks. Academic benchmarks such as tau-bench popularized this pattern for customer-service agents; use it with care, since simulated users are themselves imperfect.

Non-determinism: run it more than once

Agents are stochastic. A single run per task tells you little. Run each task several times (for example 3–5) and report:

  • Success rate per task and overall.
  • pass^k-style consistency: the share of tasks that succeed on all k runs (a stricter reliability measure that matters for customer-facing agents).
  • Variance in steps and cost.

Hands-on: an agent eval with Inspect

A minimal Inspect task with a tool and a scorer (check the Inspect documentation for current APIs):

# agent_eval.py: run with: inspect eval agent_eval.py --model <provider>/<model-id>
from inspect_ai import Task, task
from inspect_ai.dataset import Sample
from inspect_ai.scorer import includes
from inspect_ai.solver import generate, system_message, use_tools
from inspect_ai.tool import tool

ORDERS = {"A-1001": {"status": "shipped", "city": "Karachi"}, "A-1002": {"status": "processing", "city": "Dubai"}}

@tool
def get_order_status():
    async def execute(order_id: str):
        """Look up an order's status.

        Args:
            order_id: The order ID, e.g. A-1001.
        """
        order = ORDERS.get(order_id)
        return f"{order_id}: {order['status']}" if order else f"{order_id}: not found"
    return execute

@task
def order_status_agent():
    return Task(
        dataset=[
            Sample(input="Where is my order A-1001?", target="shipped"),
            Sample(input="Status of A-1002 please", target="processing"),
            Sample(input="What about order Z-9?", target="not found"),
        ],
        solver=[system_message("You are a support agent. Use tools to look up orders; never guess."),
                use_tools(get_order_status()), generate()],
        scorer=includes(),
    )

Inspect's log viewer shows each sample's full trajectory, including tool calls, which makes trajectory review practical. Add custom scorers that inspect the message history for required or forbidden tool calls.

Trajectory checks in plain Python

If you log agent runs as lists of events, trajectory rules are simple functions:

def verified_before_disclosure(events: list[dict]) -> bool:
    verified = False
    for e in events:
        if e["type"] == "tool_call" and e["name"] == "verify_customer" and e.get("result") == "ok":
            verified = True
        if e["type"] == "tool_call" and e["name"] == "get_order_details" and not verified:
            return False
    return True

def no_refund_over_limit_without_approval(events, limit_minor=500_000):
    for e in events:
        if e["type"] == "tool_call" and e["name"] == "create_refund":
            if e["args"]["amount_minor"] > limit_minor and not e["args"].get("approval_id"):
                return False
    return True

Worked example

A travel agency in Jeddah evaluated a booking-change agent on 60 scenarios, five runs each, in a sandbox with a mock airline API. Success was high on average, but consistency across all five runs was noticeably lower: on some scenarios the agent occasionally skipped fare-difference confirmation. That trajectory failure became a blocking check, and a prompt and tool-design change (making confirmation a required tool step) fixed it.

Pitfalls

  • Trusting the agent's final message instead of checking the end state.
  • One run per task.
  • Evaluating against live systems, causing side effects and irreproducible results.
  • Ignoring cost and step counts until the bill arrives.

How to measure success

For each agent release: outcome success rate, all-runs consistency, trajectory rule violations (target zero for blocking rules), and median cost and steps per successful task.

Video lecture: Evaluating agents: outcomes, trajectories and efficiency

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

  1. Evaluating agents and trajectories
  2. Analogy: judging a delivery driver
  3. 1. Outcome
  4. 2. Trajectory
  5. 3. Efficiency
  6. Sandboxes
  7. Run it more than once
  8. Hands-on
  9. Case: Jeddah booking-change agent
  10. Example: order before verification
  11. Scenario: two designs, 40 × 5 runs (illustrative)
  12. Common mistakes
  13. Deeper: cost per success compounds
  14. Watch me do it: an Inspect agent eval
  15. Recap

Lecture transcript

Evaluating agents and trajectories

A chatbot answers once. An agent plans, calls tools, reads results and decides what to do next, sometimes over dozens of steps. Two runs of the same task can take different paths and both succeed, or look identical until step six and then diverge. In this lecture you will learn to evaluate agents on outcomes, trajectories and efficiency, in safe sandboxes, with enough repeated runs to measure reliability.

Analogy: judging a delivery driver

An analogy for agent evaluation. Imagine judging a delivery driver. Did the parcel arrive? That is the outcome. But you also care how: did they run red lights, leave the van unlocked, or drive three hundred kilometers for a ten-kilometer trip? A driver who arrives by breaking the rules is a liability, and one who takes the scenic route every day is expensive. Agents are the same: outcome, trajectory and efficiency all matter.

1. Outcome

Level one is the outcome. Did the task actually get done? Check the end state of the world, not what the agent says it did. Was the refund created with the right amount? Did the file change? Does the test pass? Is the database row correct? Outcome checks are usually code, and they are the most important thing you measure. An agent that says done is not evidence.

2. Trajectory

Level two is the trajectory, the path it took. Were required steps taken, like verifying identity before revealing order details? Were forbidden actions avoided, like issuing a refund over the limit without approval? Were tool calls well formed? Did it recover sensibly from tool errors? Did it ask for clarification instead of guessing? Many serious failures happen on a path that still ends in a successful-looking outcome.

3. Efficiency

Level three is efficiency: steps, tool calls, tokens, time and cost per successful task. Two agents with the same success rate can differ several times over in cost. If you do not measure efficiency during evaluation, you will discover it on your invoice.

Sandboxes

Agents act on the world, so evaluate them in sandboxes that reset between runs. Mocked tools returning fixtures are fast and deterministic. Stateful simulators of your APIs give more realism. Containerized environments suit coding and computer-use agents; the open-source Inspect framework from the UK AI Security Institute supports Docker-based sandboxes for exactly this. And for multi-turn tasks, a simulated user, which is a model playing the customer from a scenario script, helps, though simulated users are imperfect themselves.

Run it more than once

Agents are stochastic, so one run per task tells you little. Run each task several times, say three to five. Report success rate, and also a stricter consistency measure: the share of tasks that succeed on every run. For a customer-facing agent, succeeding four times out of five means one customer in five has a bad experience on that task. Consistency is what reliability feels like to users.

Hands-on

The lesson includes a minimal Inspect task: a tool that looks up order status, three samples including an order that does not exist, and a scorer. Inspect's log viewer shows each sample's full trajectory, including tool calls, which makes reviewing paths practical. There are also plain Python trajectory rules, like verified before disclosure, and no refund over the limit without an approval ID. They are simple functions over your logged events, and they make great blocking checks.

Case: Jeddah booking-change agent

A travel agency in Jeddah evaluated a booking-change agent on sixty scenarios, five runs each, in a sandbox with a mock airline API. Average success looked strong. But all-runs consistency was noticeably lower: on some scenarios the agent occasionally skipped confirming the fare difference with the customer. That became a blocking trajectory check. The fix was partly prompt and partly tool design, making confirmation a required tool step the agent could not skip.

Example: order before verification

A simple example of a trajectory rule. The rule: never reveal order details before the customer is verified. You log an agent run as events: user message, tool call get order details, tool call verify customer, reply. The final answer looks polite and correct. But the rule function walks the events in order and sees get order details before verify customer. Fail. The outcome looked fine; the path leaked data. That is exactly why trajectory checks exist.

Scenario: two designs, 40 × 5 runs (illustrative)

Now a realistic scenario with illustrative numbers. A logistics company in Karachi compares two agent designs for rescheduling deliveries on forty scenarios, five runs each. Design A succeeds on ninety percent of runs with a median of seven tool calls. Design B also succeeds on about ninety percent, but uses a median of nineteen calls and succeeds on every run for fewer scenarios. The team ships design A: same headline success, far cheaper, and more consistent. The headline number alone would have called it a tie.

Common mistakes

Common mistakes with agent evals. Reading the agent's final message and calling it success. Running each scenario once and trusting the result. Testing against live systems, which creates real side effects and makes runs impossible to reproduce. And ignoring step counts until the monthly bill arrives. Try this now: take one scenario your agent handles every day, run it five times in a sandbox, and write down how many runs fully succeeded and how many tool calls each one used.

Deeper: cost per success compounds

One level deeper on efficiency. For the Karachi rescheduling agents, the team computed cost per successful task by adding token costs for all five runs and dividing by successful runs. Design B's extra tool calls also pulled more text into context each step, so its cost per success was several times higher, not just a little. Tool calls compound.

Watch me do it: an Inspect agent eval

Watch me do it with the Inspect task from the lesson. The tool, get order status, looks up a small dictionary of two orders: A one thousand and one shipped to Karachi, and A one thousand and two processing in Dubai. The dataset has three samples: where is my order A one thousand and one, status of A one thousand and two please, and what about order Z nine, which does not exist. The solver has a system message telling the agent to use tools and never guess, then enables the tool, then generates. The scorer checks whether the target word appears. I run inspect eval from the command line with my provider and model. Results: all three pass. But passing is not enough, so I open the log viewer and read the trajectories. For sample one, the agent called the tool once with the right ID, then answered. For sample three, it called the tool, got not found, and said so, instead of inventing a status. Good. Now I make it harder. I add a sample where the customer gives the order ID with a lowercase a and a space. The agent guesses the status without calling the tool. That is a trajectory failure the includes scorer could miss by luck. So I add a custom scorer that fails any sample where no tool call happened, and I run each sample five times to measure consistency.

Recap

Recap. Evaluate agents on outcome by checking end state, trajectory with rules for required and forbidden actions, and efficiency. Use sandboxes that reset, run tasks several times, and track all-runs consistency. Avoid trusting the agent's final message, testing against live systems, and ignoring cost. Your next step: write two trajectory rules for your agent's most dangerous actions and run each test scenario five times.

Key takeaways

  • Check outcomes by inspecting end state, not the agent's claims.
  • Encode trajectory rules for required steps and forbidden actions as blocking checks.
  • Evaluate in resettable sandboxes (mocks, simulators, containers such as Inspect's Docker sandboxes).
  • Run tasks several times; report success, all-runs consistency, steps and cost.

Try it

Write two trajectory rules for your agent's riskiest actions and run five repetitions of each test scenario in a sandbox, reporting consistency.