---
title: "Evaluating agents: task success, trajectories and judges"
description: "Why agent evaluation is different A chatbot answer can be scored once. An agent produces a trajectory : a sequence of decisions, tool calls and…"
url: https://optimizeall.com/learn/ai-agents-engineering/evaluating-agents
updated: 2026-10-05
---

Building Production AI Agents · Evaluation and observability · lesson 12 of 18 · 17 min

# Evaluating agents: task success, trajectories and judges

## Why agent evaluation is different

A chatbot answer can be scored once. An agent produces a **trajectory**: a sequence of decisions, tool calls and intermediate states that ends in an outcome. Two agents can reach the same answer, one in three calls and one in thirty with a risky action along the way. Agents are also non-deterministic: the same task can succeed on one run and fail on the next. Your evaluation must capture outcome, process, cost and consistency.

## The four layers of agent evals

1. **Outcome**: did the task succeed? Prefer **programmatic checks** against the final environment state (the ticket is closed with the right tag; the spreadsheet cell has the correct total; the tests pass) over judging text.
2. **Trajectory**: was the path acceptable? Right tools, valid arguments, no forbidden actions, reasonable number of steps, approvals requested when required.
3. **Quality of outputs**: for text deliverables (memos, emails), use rubrics scored by humans or an **LLM judge** calibrated against humans.
4. **Efficiency**: tokens, cost, latency and tool calls per task.

## Building the eval set

- Start with **20–50 real tasks** from logs or user interviews, not synthetic toy tasks. Include easy, typical and hard cases, and adversarial ones (ambiguous inputs, missing data, injection attempts).
- For each task define: input, environment setup (fixtures, mock tools or sandbox data), success criteria (code checks where possible), forbidden actions, and a reference trajectory if useful.
- Version the set; grow it every time production reveals a new failure. A bug report becomes a test case.

## Consistency: pass@k versus pass^k

Run each task several times. **pass@k** (at least one of k runs succeeds) measures capability. **pass^k** (all k runs succeed), popularized by the tau-bench agent benchmark, measures reliability, which is what customers experience. An agent with 80% single-run success has pass^3 of roughly 51% if runs are independent (0.8³). Report both.

## LLM-as-judge done properly

- Give the judge a **specific rubric** with examples of each score, and ask for a short justification before the score.
- Use **pairwise comparison** (A vs B) for subtle quality differences; it is more reliable than absolute scores.
- **Calibrate**: have humans label 50–100 items and check agreement with the judge before trusting it; re-check when you change the judge model.
- Watch for known biases (position bias, preference for longer answers, self-preference for outputs from the same model family). Randomize order and control length.

## Trajectory checks you can code

- Tool sequence contains required steps (e.g., `search_crm` before `draft_email`).
- No tier-3 tool executed without an approval event.
- Step count within a threshold.
- No tool errors left unhandled.
- No calls to tools outside the allowed set (a strong injection signal).

## Worked example: evaluating a Pakistani fintech's KYC-ops agent

The agent reviews onboarding cases, requests missing documents and flags risky cases for analysts.

- 40 anonymized historical cases with known correct outcomes (approve / request docs / escalate).
- Outcome check: decision matches label; requested documents match the missing list.
- Trajectory check: must call `check_sanctions_list` before any approve decision; must never call `approve_account` (humans approve).
- Judge: rubric for the clarity and politeness of document-request messages in English and Urdu, calibrated against two analysts.
- Runs: 3 per case; report pass@1, pass^3, cost per case, and per-category failure breakdown.

The first run revealed the agent skipped the sanctions check when the case "looked obviously fine" — a trajectory failure an outcome-only eval would have missed.

## Hands-on: a minimal agent eval harness

```python
import json, statistics
from dataclasses import dataclass, field

@dataclass
class Case:
    id: str
    goal: str
    check: callable                       # (final_state, trace) -> bool
    must_call_before: tuple = ()           # e.g., ("check_sanctions_list", "decide")
    forbidden: set = field(default_factory=set)

def trajectory_ok(trace: list[dict], case: Case) -> list[str]:
    tools = [t["tool"] for t in trace if t["type"] == "tool_call"]
    problems = [f"forbidden tool {t}" for t in tools if t in case.forbidden]
    if case.must_call_before:
        a, b = case.must_call_before
        if b in tools and (a not in tools or tools.index(a) > tools.index(b)):
            problems.append(f"{a} must precede {b}")
    return problems

def run_suite(cases: list[Case], run_agent, k: int = 3) -> dict:
    rows = []
    for c in cases:
        results = []
        for _ in range(k):
            final_state, trace, usage = run_agent(c.goal)   # your agent returns these
            traj = trajectory_ok(trace, c)
            results.append({"ok": c.check(final_state, trace) and not traj, "traj": traj,
                            "cost": usage["cost_usd"], "steps": len(trace)})
        rows.append({"id": c.id, "pass_at_1": results[0]["ok"],
                     "pass_all_k": all(r["ok"] for r in results),
                     "cost": statistics.mean(r["cost"] for r in results),
                     "problems": sorted({p for r in results for p in r["traj"]})})
    summary = {"pass_at_1": sum(r["pass_at_1"] for r in rows) / len(rows),
               f"pass^{k}": sum(r["pass_all_k"] for r in rows) / len(rows),
               "mean_cost_usd": statistics.mean(r["cost"] for r in rows)}
    print(json.dumps(summary, indent=2))
    return {"summary": summary, "rows": rows}
```

Wire it into CI so any change to prompts, tools or model versions runs the suite and blocks merges that reduce pass^k or raise cost beyond a threshold. Hosted tools (for example LangSmith, Braintrust, Langfuse, provider eval dashboards) add dataset management and comparison views; the logic stays the same.

## Pitfalls

- Evaluating only final text with a vague judge prompt.
- Tiny eval sets that make a 5% change meaningless noise.
- Evaluating on live systems with real side effects; use sandboxes and fixtures.
- Never refreshing the set, so it drifts from real traffic.

## Measuring success

Your eval program itself has metrics: coverage of production task types, judge–human agreement, time to run the suite, and how many production incidents were caught first by evals.

## Video lecture: Evaluating agents: task success, trajectories and judges

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. Evaluating agents
2. Why it matters
3. Four layers of agent evals
4. Build the eval set
5. Consistency
6. Simple example: room booking
7. LLM-as-judge
8. Example: fintech onboarding agent
9. Hands-on: eval harness
10. Pitfalls
11. How many cases?
12. Deeper: the fintech KYC eval (illustrative)
13. Watch me do it: run_suite()
14. Try this now
15. Recap

## Lecture transcript

### Evaluating agents

How do you know your agent is good? Not because the demo looked great, and not because one run worked. Agents take different paths each time, and a right answer can hide a dangerous process. In this lesson you'll learn how professionals evaluate agents: outcomes, trajectories, judges and consistency.

### Why it matters

Why does evaluation matter so much for agents in particular? Because agents fail quietly. A chatbot's bad answer is visible. An agent can take the right final action through a dangerous path, or succeed on Monday and fail on Tuesday with the same input. Think of a driving test. The examiner doesn't only check that you arrived. They watch whether you signaled, checked mirrors and stopped at the red light. Trajectory evaluation is watching the driving, not just the destination.

### Four layers of agent evals

A chatbot answer can be scored once. An agent produces a trajectory, a sequence of decisions and tool calls that ends in an outcome. So measure four layers. Outcome: did it succeed? Trajectory: was the path acceptable? Output quality: for things like memos and emails, is the writing good? And efficiency: tokens, cost, time and tool calls. For outcomes, prefer code that checks the final state of the world, like the ticket is closed with the right tag, or the tests pass, over asking a model whether the text sounds right.

### Build the eval set

Build your eval set from reality. Start with twenty to fifty real tasks from logs or interviews, covering easy, typical, hard and adversarial cases, like ambiguous inputs, missing data and injection attempts. For each task, define the input, the environment setup with fixtures or sandbox data, the success check, any forbidden actions, and optionally a reference path. Version the set, and every time production shows a new failure, turn that bug report into a test case.

### Consistency

Agents are non deterministic, so run each task several times. Two numbers matter. Pass at k means at least one of k runs succeeded. That measures capability. Pass hat k, written with a caret, means all k runs succeeded. That measures reliability, which is what customers actually feel. Here's why it matters: an agent that succeeds eighty percent of the time on a single run has a pass hat three of only about fifty one percent, if runs are independent. Report both.

### Simple example: room booking

A simple example. Your agent books meeting rooms. Test case: book a room for six people on Thursday at ten, with a screen. The outcome check looks at the calendar system afterwards: is there exactly one booking, Thursday at ten, in a room with capacity six or more and a screen? The trajectory check confirms it searched availability before booking and didn't cancel anyone else's meeting. Run it three times. If it passes twice and fails once, your single run success looks fine, but your reliability needs work.

### LLM-as-judge

For text outputs you'll often use a model as the judge. Do it properly. Give it a specific rubric with examples of each score, and ask for a short justification before the score. Use pairwise comparisons, A versus B, for subtle differences. Calibrate: have humans label fifty to a hundred items and check agreement before you trust the judge. And watch for known biases like position bias, a preference for longer answers, and a preference for outputs from its own model family. Randomize order and control length.

### Example: fintech onboarding agent

A worked example. A fintech in Pakistan uses an agent to review customer onboarding cases, request missing documents and flag risky cases. The team uses forty anonymized historical cases with known outcomes. Outcome checks compare decisions and requested documents. Trajectory checks require a sanctions list check before any approval, and forbid the agent from approving accounts at all. A calibrated judge scores document request messages in English and Urdu. The first run exposed that the agent skipped the sanctions check when a case looked obviously fine. An outcome only eval would have missed it completely.

### Hands-on: eval harness

The hands on harness in the lesson runs each case k times, checks the final state with your function, checks the trajectory for forbidden tools and required ordering, and reports pass at one, pass hat k and mean cost. Wire it into continuous integration so any change to prompts, tools or model versions runs the suite automatically, and block merges that reduce reliability or raise cost past a threshold. Hosted tools add dashboards and dataset management, but the logic is the same.

### Pitfalls

Common pitfalls: judging only final text with a vague prompt, eval sets so small that a five percent change is just noise, evaluating on live systems with real side effects instead of sandboxes, and never refreshing the set. Your eval program has its own metrics too: coverage of real task types, judge to human agreement, how long the suite takes, and how many incidents your evals caught before production did.

### How many cases?

A practical question: how many test cases do you need? Start with twenty to fifty real tasks, run three times each. That's enough to catch big problems and compare two versions with obvious differences. As the agent matures and changes become subtler, grow the set toward a hundred or more, especially covering your most common and most costly task types. And weight your attention by business impact: a failure on refunds matters more than a failure on a greeting.

### Deeper: the fintech KYC eval (illustrative)

Let's deepen the Pakistani fintech example. Forty historical onboarding cases, three runs each. The first evaluation showed pass at one around three quarters and pass hat three much lower, illustrative numbers, and the trajectory checks explained why: in about one run in six, the agent skipped the sanctions check on cases that looked obviously fine. The team made the sanctions tool the only way to unlock a decision, by requiring its result id in the decision tool's arguments. That's mistake proofing again. The rerun showed zero skipped checks, and pass hat three rose substantially. The compliance officer asked for one thing: keep those forty cases as a permanent regression suite, rerun on every change.

### Watch me do it: run_suite()

Watch me do it. Let's walk through the eval harness. A case has an id, a goal, a check function that receives the final state and trace, an optional pair of tools that must be called in order, and a set of forbidden tools. Trajectory ok pulls the tool names from the trace, flags any forbidden ones, and if the must call before pair is set, flags the case when the second tool appears without the first before it. Run suite loops over cases and runs the agent k times. Each run records whether the check passed with no trajectory problems, the cost and the number of steps. Per case, it stores pass at one from the first run, and pass all k if every run passed. The summary prints pass at one, pass hat k and mean cost. I'll run it with k equals three on ten cases: pass at one point eight, pass hat three point six, so reliability is my next job.

### Try this now

Try this now. Write ten test cases for your agent. For each, define what the world should look like afterwards, and write a small function that checks it. Add at least one trajectory rule per case, like must call this tool first, or must never call that one. Run each case three times and calculate pass at one and pass hat three. The gap between those two numbers tells you whether your next job is capability or reliability.

### Recap

To recap: evaluate outcomes, trajectories, quality and efficiency; build sets from real tasks; report pass at k and pass hat k; calibrate judges; and run it all in CI. Your next step: write ten eval cases for your agent, each with a programmatic check and at least one trajectory rule. Run each three times and report your numbers.

## Key takeaways

- Agent evals cover outcome, trajectory, output quality and efficiency.
- Prefer programmatic checks of final environment state over judging text.
- Report pass@k for capability and pass^k for reliability across repeated runs.
- Calibrate LLM judges against human labels and control for known biases.
- Run the suite in CI on every prompt, tool or model change.

## Try it

Write ten eval cases for your agent with a programmatic success check and at least one trajectory rule each. Run each three times and report pass@1 and pass^3.

- [Previous: Guardrails, permissions and prompt-injection defense](https://optimizeall.com/learn/ai-agents-engineering/guardrails-permissions-prompt-injection)
- [Next: Observability: tracing, logging and dashboards for agents](https://optimizeall.com/learn/ai-agents-engineering/observability-and-tracing)
- [All lessons of Building Production AI Agents](https://optimizeall.com/learn/ai-agents-engineering)
