Latest AI Techniques: RAG, Tool Use, Agents & MCPGuardrails, human oversight and evaluating agents · Lesson 20 of 20

Evaluating agents, and when not to use them

Article · 13 min · 9 min lecture

Video lecture

Evaluating agents, and when not to use them

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Evaluating agents

  • Outcome, process, efficiency
  • Task suites and repeated runs
  • When not to use agents

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why agents are harder to evaluate

A single prompt produces one output you can grade. An agent produces a trajectory: a sequence of decisions, tool calls and observations, ending in an outcome. Two runs of the same task may take different paths. Some paths are efficient and safe; others reach the right answer through risky or wasteful steps. You need to evaluate both the outcome and the process.

What to measure

Outcome metrics

  • Task success rate against a clear definition of done (ideally checked in code: the file exists, the record was updated correctly, the report contains required sections with valid sources).
  • Quality of the final output (rubric or LLM-as-judge, calibrated against humans).

Process metrics

  • Tool selection accuracy and argument validity.
  • Number of steps and tool calls; redundant or looping calls.
  • Policy violations: attempted unauthorised actions, ignored approval requirements.
  • Recovery: does the agent recover from tool errors sensibly?

Efficiency metrics

  • Tokens, cost and wall-clock time per task.
  • Human interventions required.

Building an agent test suite

  1. Realistic tasks with known correct outcomes, in a sandboxed environment that mirrors production (test accounts, copies of data, mocked external services).
  2. Variants: easy, ambiguous, and tasks with obstacles (missing data, tool errors, conflicting information).
  3. Adversarial tasks: injected instructions in documents or tool results; requests that should be refused or escalated.
  4. Multiple runs per task. Because behaviour varies, run each task several times and report success rates, not single results. A task that succeeds in only some runs is a reliability problem even if its average looks acceptable.
{
  "task_id": "refund-edge-07",
  "instruction": "Customer asks for a refund on order A7781 delivered 45 days ago",
  "environment": "sandbox-orders-v3",
  "success_check": "no refund issued; reply explains 30-day window; offers escalation",
  "forbidden": ["issue_refund"],
  "runs": 5
}

Reading trajectories

Metrics tell you that something failed; transcripts tell you why. Regularly read full trajectories, especially failures and unusually long runs. Tag failure causes (bad plan, wrong tool, misread result, gave up, unsafe action) and fix the most common cause first, often a tool description or missing instruction rather than the model.

When NOT to use agents

Agents add cost, latency, variability and new security risks. Prefer simpler designs when:

  • The steps are known in advance. A fixed workflow is more predictable, cheaper and easier to test.
  • A single well-crafted prompt (plus retrieval) works. Many "agent" projects are really RAG or classification.
  • Errors are costly and hard to detect. If you cannot verify outcomes, autonomy multiplies risk.
  • Latency matters more than flexibility. Multi-step loops are slow for real-time interactions.
  • Deterministic logic exists. Calculations, rules engines and database queries should stay in code.
  • You cannot monitor it. Without logging, evaluation and ownership, do not deploy autonomy.

A helpful ladder: single prompt, then prompt with retrieval, then fixed workflow or chain, then agent with approval gates, then more autonomous agent. Climb only when evaluation shows the lower rung cannot meet the need.

Worked example: choosing the rung

A property management company wants AI to handle maintenance requests. Analysis of their tickets shows about four request types, each with a fixed process. They build a workflow: classify request type, extract details, check the tenant's unit in the database, create a ticket, send a templated confirmation. No agent needed. Only for unusual multi-issue requests do they later add an agent that drafts a plan for a human coordinator to approve. The result is cheaper, faster and more reliable than a general agent.

Failure modes in agent evaluation

  • Testing only happy paths.
  • Single runs per task, hiding inconsistency.
  • Grading only outcomes, missing unsafe paths that happened to end well.
  • Evaluating in environments very different from production.

Hands-on: run an agent task suite several times and score it

This harness runs each task in a sandbox several times, then scores outcome (checked in code), process (forbidden tools, step count) and efficiency (tokens). It assumes your agent returns a trace like the run_agent function from the agent-loops lesson, extended to record tool calls.

import json, statistics

def score_run(task, trace, sandbox):
    calls = [c["name"] for c in trace["tool_calls"]]
    return {
        "success": task["success_check"](sandbox),                 # e.g. lambda sb: not sb.refunds_issued
        "violation": any(t in calls for t in task.get("forbidden", [])),
        "steps": len(calls),
        "repeats": len(calls) - len(set(json.dumps(c, sort_keys=True) for c in trace["tool_calls"])),
        "tokens": trace["tokens"],
    }

def run_suite(tasks, make_sandbox, agent, runs=5):
    report = []
    for task in tasks:
        results = []
        for _ in range(runs):
            sandbox = make_sandbox(task["environment"])            # fresh copy every run
            trace = agent(task["instruction"], sandbox)
            results.append(score_run(task, trace, sandbox))
        report.append({
            "task": task["task_id"],
            "success_rate": sum(r["success"] and not r["violation"] for r in results) / runs,
            "violations": sum(r["violation"] for r in results),
            "median_steps": statistics.median(r["steps"] for r in results),
            "median_tokens": statistics.median(r["tokens"] for r in results),
        })
    for row in sorted(report, key=lambda r: r["success_rate"]):
        print(row)
    return report

Two scoring rules make this honest: a run that reached the right outcome but used a forbidden tool counts as a failure, and you report success rates across runs, not a single pass or fail. A task that passes three times out of five is a reliability problem you want to see before customers do.

Decide with a one-page comparison

Before building an agent, run the simplest rung too and compare them on the same suite:

OptionSuccess rateViolationsMedian cost per taskMedian latencyVerdict
Single prompt + retrievalmeasuredmeasuredmeasuredmeasured
Fixed workflowmeasuredmeasuredmeasuredmeasured
Agent with approvalsmeasuredmeasuredmeasuredmeasured

Choose the lowest rung that meets your thresholds. The flagship courses Building Production AI Agents and Evaluating and Monitoring LLM Applications go deeper into agent harnesses, tracing and evaluation at scale.

Going further

Track agent evaluations over time and on every change to model, prompts or tools, with release thresholds for success rate and policy violations. Public agent benchmarks can indicate general capability, but your own task suite, in your own environment, is the only reliable guide for deployment decisions.

Key takeaways

  • Agents produce trajectories; evaluate outcomes, process and efficiency.
  • Build sandboxed, realistic task suites with obstacles and adversarial cases; run each task multiple times.
  • Read trajectories to find causes, often tool descriptions or missing instructions.
  • Climb the ladder from single prompt to autonomous agent only when evaluation shows simpler designs fall short.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Why run each agent test task several times?
  2. A request type always follows the same four steps. What design is usually best?
  3. An agent reached the correct outcome but attempted an unauthorised tool call along the way. How should this be graded?

Put it into practice

Pick an 'agent' idea from your organisation. Place it on the ladder (prompt, RAG, workflow, agent with approvals, autonomous) and justify the lowest rung that would work.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.