Building Production AI AgentsEvaluation and observability · Lesson 12 of 18

Evaluating agents: task success, trajectories and judges

Article · 17 min · 9 min lecture

Video lecture

Evaluating agents: task success, trajectories and judges

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Evaluating agents

  • Outcomes and trajectories
  • Consistency: pass@k and pass^k
  • LLM judges done properly

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why agent evaluation is different

A chatbot answer can be scored once. An agent produces a trajectory: a sequence of decisions, tool calls and intermediate states that ends in an outcome. Two agents can reach the same answer, one in three calls and one in thirty with a risky action along the way. Agents are also non-deterministic: the same task can succeed on one run and fail on the next. Your evaluation must capture outcome, process, cost and consistency.

The four layers of agent evals

  1. Outcome: did the task succeed? Prefer programmatic checks against the final environment state (the ticket is closed with the right tag; the spreadsheet cell has the correct total; the tests pass) over judging text.
  2. Trajectory: was the path acceptable? Right tools, valid arguments, no forbidden actions, reasonable number of steps, approvals requested when required.
  3. Quality of outputs: for text deliverables (memos, emails), use rubrics scored by humans or an LLM judge calibrated against humans.
  4. Efficiency: tokens, cost, latency and tool calls per task.

Building the eval set

  • Start with 20–50 real tasks from logs or user interviews, not synthetic toy tasks. Include easy, typical and hard cases, and adversarial ones (ambiguous inputs, missing data, injection attempts).
  • For each task define: input, environment setup (fixtures, mock tools or sandbox data), success criteria (code checks where possible), forbidden actions, and a reference trajectory if useful.
  • Version the set; grow it every time production reveals a new failure. A bug report becomes a test case.

Consistency: pass@k versus pass^k

Run each task several times. pass@k (at least one of k runs succeeds) measures capability. pass^k (all k runs succeed), popularized by the tau-bench agent benchmark, measures reliability, which is what customers experience. An agent with 80% single-run success has pass^3 of roughly 51% if runs are independent (0.8³). Report both.

LLM-as-judge done properly

  • Give the judge a specific rubric with examples of each score, and ask for a short justification before the score.
  • Use pairwise comparison (A vs B) for subtle quality differences; it is more reliable than absolute scores.
  • Calibrate: have humans label 50–100 items and check agreement with the judge before trusting it; re-check when you change the judge model.
  • Watch for known biases (position bias, preference for longer answers, self-preference for outputs from the same model family). Randomize order and control length.

Trajectory checks you can code

  • Tool sequence contains required steps (e.g., search_crm before draft_email).
  • No tier-3 tool executed without an approval event.
  • Step count within a threshold.
  • No tool errors left unhandled.
  • No calls to tools outside the allowed set (a strong injection signal).

Worked example: evaluating a Pakistani fintech's KYC-ops agent

The agent reviews onboarding cases, requests missing documents and flags risky cases for analysts.

  • 40 anonymized historical cases with known correct outcomes (approve / request docs / escalate).
  • Outcome check: decision matches label; requested documents match the missing list.
  • Trajectory check: must call check_sanctions_list before any approve decision; must never call approve_account (humans approve).
  • Judge: rubric for the clarity and politeness of document-request messages in English and Urdu, calibrated against two analysts.
  • Runs: 3 per case; report pass@1, pass^3, cost per case, and per-category failure breakdown.

The first run revealed the agent skipped the sanctions check when the case "looked obviously fine" — a trajectory failure an outcome-only eval would have missed.

Hands-on: a minimal agent eval harness

import json, statistics
from dataclasses import dataclass, field

@dataclass
class Case:
    id: str
    goal: str
    check: callable                       # (final_state, trace) -> bool
    must_call_before: tuple = ()           # e.g., ("check_sanctions_list", "decide")
    forbidden: set = field(default_factory=set)

def trajectory_ok(trace: list[dict], case: Case) -> list[str]:
    tools = [t["tool"] for t in trace if t["type"] == "tool_call"]
    problems = [f"forbidden tool {t}" for t in tools if t in case.forbidden]
    if case.must_call_before:
        a, b = case.must_call_before
        if b in tools and (a not in tools or tools.index(a) > tools.index(b)):
            problems.append(f"{a} must precede {b}")
    return problems

def run_suite(cases: list[Case], run_agent, k: int = 3) -> dict:
    rows = []
    for c in cases:
        results = []
        for _ in range(k):
            final_state, trace, usage = run_agent(c.goal)   # your agent returns these
            traj = trajectory_ok(trace, c)
            results.append({"ok": c.check(final_state, trace) and not traj, "traj": traj,
                            "cost": usage["cost_usd"], "steps": len(trace)})
        rows.append({"id": c.id, "pass_at_1": results[0]["ok"],
                     "pass_all_k": all(r["ok"] for r in results),
                     "cost": statistics.mean(r["cost"] for r in results),
                     "problems": sorted({p for r in results for p in r["traj"]})})
    summary = {"pass_at_1": sum(r["pass_at_1"] for r in rows) / len(rows),
               f"pass^{k}": sum(r["pass_all_k"] for r in rows) / len(rows),
               "mean_cost_usd": statistics.mean(r["cost"] for r in rows)}
    print(json.dumps(summary, indent=2))
    return {"summary": summary, "rows": rows}

Wire it into CI so any change to prompts, tools or model versions runs the suite and blocks merges that reduce pass^k or raise cost beyond a threshold. Hosted tools (for example LangSmith, Braintrust, Langfuse, provider eval dashboards) add dataset management and comparison views; the logic stays the same.

Pitfalls

  • Evaluating only final text with a vague judge prompt.
  • Tiny eval sets that make a 5% change meaningless noise.
  • Evaluating on live systems with real side effects; use sandboxes and fixtures.
  • Never refreshing the set, so it drifts from real traffic.

Measuring success

Your eval program itself has metrics: coverage of production task types, judge–human agreement, time to run the suite, and how many production incidents were caught first by evals.

Key takeaways

  • Agent evals cover outcome, trajectory, output quality and efficiency.
  • Prefer programmatic checks of final environment state over judging text.
  • Report pass@k for capability and pass^k for reliability across repeated runs.
  • Calibrate LLM judges against human labels and control for known biases.
  • Run the suite in CI on every prompt, tool or model change.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. An agent reaches the correct KYC decision but skipped the mandatory sanctions check. Which eval layer catches this?
  2. An agent succeeds on 80% of single runs. Assuming independence, roughly what is pass^3?
  3. What should you do before trusting an LLM judge's scores?

Put it into practice

Write ten eval cases for your agent with a programmatic success check and at least one trajectory rule each. Run each three times and report pass@1 and pass^3.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.