---
title: "Evaluating agents, and when not to use them"
description: "Why agents are harder to evaluate A single prompt produces one output you can grade. An agent produces a trajectory : a sequence of decisions, tool calls…"
url: https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp/evaluating-agents-and-when-not-to
updated: 2026-10-05
---

Latest AI Techniques: RAG, Tool Use, Agents & MCP · Guardrails, human oversight and evaluating agents · lesson 20 of 20 · 13 min

# Evaluating agents, and when not to use them

## Why agents are harder to evaluate

A single prompt produces one output you can grade. An agent produces a **trajectory**: a sequence of decisions, tool calls and observations, ending in an outcome. Two runs of the same task may take different paths. Some paths are efficient and safe; others reach the right answer through risky or wasteful steps. You need to evaluate both the **outcome** and the **process**.

## What to measure

**Outcome metrics**

- Task success rate against a clear definition of done (ideally checked in code: the file exists, the record was updated correctly, the report contains required sections with valid sources).
- Quality of the final output (rubric or LLM-as-judge, calibrated against humans).

**Process metrics**

- Tool selection accuracy and argument validity.
- Number of steps and tool calls; redundant or looping calls.
- Policy violations: attempted unauthorised actions, ignored approval requirements.
- Recovery: does the agent recover from tool errors sensibly?

**Efficiency metrics**

- Tokens, cost and wall-clock time per task.
- Human interventions required.

## Building an agent test suite

1. **Realistic tasks** with known correct outcomes, in a sandboxed environment that mirrors production (test accounts, copies of data, mocked external services).
2. **Variants:** easy, ambiguous, and tasks with obstacles (missing data, tool errors, conflicting information).
3. **Adversarial tasks:** injected instructions in documents or tool results; requests that should be refused or escalated.
4. **Multiple runs per task.** Because behaviour varies, run each task several times and report success rates, not single results. A task that succeeds in only some runs is a reliability problem even if its average looks acceptable.

```json
{
  "task_id": "refund-edge-07",
  "instruction": "Customer asks for a refund on order A7781 delivered 45 days ago",
  "environment": "sandbox-orders-v3",
  "success_check": "no refund issued; reply explains 30-day window; offers escalation",
  "forbidden": ["issue_refund"],
  "runs": 5
}
```

## Reading trajectories

Metrics tell you *that* something failed; transcripts tell you *why*. Regularly read full trajectories, especially failures and unusually long runs. Tag failure causes (bad plan, wrong tool, misread result, gave up, unsafe action) and fix the most common cause first, often a tool description or missing instruction rather than the model.

## When NOT to use agents

Agents add cost, latency, variability and new security risks. Prefer simpler designs when:

- **The steps are known in advance.** A fixed workflow is more predictable, cheaper and easier to test.
- **A single well-crafted prompt (plus retrieval) works.** Many "agent" projects are really RAG or classification.
- **Errors are costly and hard to detect.** If you cannot verify outcomes, autonomy multiplies risk.
- **Latency matters more than flexibility.** Multi-step loops are slow for real-time interactions.
- **Deterministic logic exists.** Calculations, rules engines and database queries should stay in code.
- **You cannot monitor it.** Without logging, evaluation and ownership, do not deploy autonomy.

A helpful ladder: single prompt, then prompt with retrieval, then fixed workflow or chain, then agent with approval gates, then more autonomous agent. Climb only when evaluation shows the lower rung cannot meet the need.

## Worked example: choosing the rung

A property management company wants AI to handle maintenance requests. Analysis of their tickets shows about four request types, each with a fixed process. They build a workflow: classify request type, extract details, check the tenant's unit in the database, create a ticket, send a templated confirmation. No agent needed. Only for unusual multi-issue requests do they later add an agent that drafts a plan for a human coordinator to approve. The result is cheaper, faster and more reliable than a general agent.

## Failure modes in agent evaluation

- Testing only happy paths.
- Single runs per task, hiding inconsistency.
- Grading only outcomes, missing unsafe paths that happened to end well.
- Evaluating in environments very different from production.

## Hands-on: run an agent task suite several times and score it

This harness runs each task in a sandbox several times, then scores outcome (checked in code), process (forbidden tools, step count) and efficiency (tokens). It assumes your agent returns a trace like the `run_agent` function from the agent-loops lesson, extended to record tool calls.

```python
import json, statistics

def score_run(task, trace, sandbox):
    calls = [c["name"] for c in trace["tool_calls"]]
    return {
        "success": task["success_check"](sandbox),                 # e.g. lambda sb: not sb.refunds_issued
        "violation": any(t in calls for t in task.get("forbidden", [])),
        "steps": len(calls),
        "repeats": len(calls) - len(set(json.dumps(c, sort_keys=True) for c in trace["tool_calls"])),
        "tokens": trace["tokens"],
    }

def run_suite(tasks, make_sandbox, agent, runs=5):
    report = []
    for task in tasks:
        results = []
        for _ in range(runs):
            sandbox = make_sandbox(task["environment"])            # fresh copy every run
            trace = agent(task["instruction"], sandbox)
            results.append(score_run(task, trace, sandbox))
        report.append({
            "task": task["task_id"],
            "success_rate": sum(r["success"] and not r["violation"] for r in results) / runs,
            "violations": sum(r["violation"] for r in results),
            "median_steps": statistics.median(r["steps"] for r in results),
            "median_tokens": statistics.median(r["tokens"] for r in results),
        })
    for row in sorted(report, key=lambda r: r["success_rate"]):
        print(row)
    return report
```

Two scoring rules make this honest: a run that reached the right outcome but used a forbidden tool counts as a failure, and you report success **rates** across runs, not a single pass or fail. A task that passes three times out of five is a reliability problem you want to see before customers do.

## Decide with a one-page comparison

Before building an agent, run the simplest rung too and compare them on the same suite:

| Option | Success rate | Violations | Median cost per task | Median latency | Verdict |
|---|---|---|---|---|---|
| Single prompt + retrieval | measured | measured | measured | measured | |
| Fixed workflow | measured | measured | measured | measured | |
| Agent with approvals | measured | measured | measured | measured | |

Choose the lowest rung that meets your thresholds. The flagship courses **Building Production AI Agents** and **Evaluating and Monitoring LLM Applications** go deeper into agent harnesses, tracing and evaluation at scale.

## Going further

Track agent evaluations over time and on every change to model, prompts or tools, with release thresholds for success rate and policy violations. Public agent benchmarks can indicate general capability, but your own task suite, in your own environment, is the only reliable guide for deployment decisions.

## Video lecture: Evaluating agents, and when not to use them

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. Evaluating agents
2. Analogy: the driving test
3. Three lenses
4. Build the suite
5. Simple example: a 45-day refund
6. Repeat and be strict
7. Read the trajectories
8. When not to use agents
9. Worked example + hands-on
10. Common mistakes
11. How you'll know it's right
12. Business example (illustrative)
13. Watch me do it: suite runner
14. Recap
15. Try this now (20 minutes)

## Lecture transcript

### Evaluating agents

Two runs of the same agent task. Both end with the right answer. One took four tidy steps. The other took nineteen, tried to call a tool it wasn't allowed to use, and got lucky. If you only grade the final answer, you'd score them the same. In this lesson you'll learn to evaluate agents on outcome, process and efficiency, build a task suite that reveals reliability problems, and, just as important, recognise when you shouldn't use an agent at all.

### Analogy: the driving test

Here's an analogy for evaluating agents. Judging a driving test only by whether the car arrived is a terrible idea. You'd pass someone who ran two red lights and got lucky. Examiners watch the whole journey: mirrors, signals, speed, safety. Agent evaluation is the same. The outcome matters, but so does the route: which tools it used, what it tried, and whether it broke any rules on the way.

### Three lenses

An agent produces a trajectory, a sequence of decisions, tool calls and observations ending in an outcome. So measure three things. Outcome: did it succeed against a clear definition of done, ideally checked in code, like the record was updated correctly or the report has all required sections with valid sources? Process: did it choose the right tools with valid arguments, avoid loops, respect policies and recover from errors? And efficiency: tokens, cost, time and human interventions per task.

### Build the suite

Build a test suite of realistic tasks with known correct outcomes, run in a sandbox that mirrors production: test accounts, copies of data and mocked external services. Include easy tasks, ambiguous ones and tasks with obstacles, like missing data, tool errors and conflicting information. Add adversarial tasks, such as injected instructions inside documents and requests that should be refused or escalated. For each task, write the success check and the list of forbidden actions.

### Simple example: a 45-day refund

A simple example of scoring. Your agent handles refund requests, and one test says: a customer asks for a refund forty-five days after delivery; the policy window is thirty days. The success check is in code: no refund was issued, and the reply mentions the thirty-day window and offers escalation. The forbidden tool is issue refund. Run it five times. Four runs pass. One run issues the refund. That's an eighty percent pass rate with a policy violation, which blocks release.

### Repeat and be strict

Run each task several times. Agent behaviour varies from run to run, so a single result can mislead. Report success rates. A task that succeeds three times in five is a reliability problem, even if the average across the suite looks fine. And grade process as well as outcome. A run that reached the right answer but attempted an unauthorised action is a failure, full stop.

### Read the trajectories

Metrics tell you that something failed. Transcripts tell you why. Read full trajectories regularly, especially failures and unusually long runs. Tag the cause, such as bad plan, wrong tool, misread result, gave up, or unsafe action, and fix the most common cause first. Very often it's a tool description or a missing instruction, not the model.

### When not to use agents

Now the discipline that saves the most money: knowing when not to use an agent. Agents add cost, latency, variability and security risk. Prefer something simpler when the steps are known in advance, when one good prompt plus retrieval works, when errors are costly and hard to detect, when latency matters more than flexibility, when deterministic logic exists, or when you can't monitor it. Climb the ladder: single prompt, then retrieval, then a fixed workflow, then an agent with approvals, then more autonomy, and only when evaluation shows the lower rung falls short.

### Worked example + hands-on

Here's a worked example. A property management company wants AI for maintenance requests. Analysis shows about four request types, each with a fixed process. So they build a workflow: classify, extract details, check the tenant's unit, create a ticket, confirm. No agent. Only for unusual multi-issue requests do they add an agent that drafts a plan for a human coordinator to approve. It's cheaper, faster and more reliable than a general agent. The lesson's hands-on harness runs tasks several times in fresh sandboxes and scores success rates, violations, steps and tokens, plus a table for comparing rungs.

### Common mistakes

Common mistakes in agent evaluation. Testing only happy paths. Running each task once. Grading only the final answer. Using a test environment that looks nothing like production. And building an agent before trying the simpler rungs, so you never learn that a workflow would have been cheaper and more reliable. The ladder isn't a formality. It's the most reliable cost-saving tool in this course.

### How you'll know it's right

How will you know you chose the right rung and evaluated it well? Your suite's success rates match what users experience in production. Policy violations are zero, not just rare. Cost per completed task is acceptable against the value. And periodically, you re-test a simpler rung. As models improve, a task that once needed an agent may now work with a single prompt and retrieval, at a fraction of the cost.

### Business example (illustrative)

More detail on the property manager, illustrative. About ninety-five percent of maintenance requests fit the four fixed types, and the workflow handles them at a few cents each with near-perfect consistency. The remaining five percent, multi-issue requests, go to the planning agent, whose plans the coordinator approves or edits. A general agent for everything, trialled first, cost several times more per request and passed the same test suite less often.

### Watch me do it: suite runner

Watch me do it. I open the suite runner. First, score run takes a task, the agent's trace and the sandbox. It checks success using the task's success check in code, flags a violation if any forbidden tool appears in the calls, counts steps, counts repeated identical calls and reads the tokens. Next, run suite loops over tasks, and for each one runs the agent five times, each time in a fresh sandbox, so runs don't contaminate each other. A run only counts as a success if it succeeded and had no violation. Then I print each task's success rate, violations, median steps and median tokens, lowest success first. I run it on six tasks. Refund edge seven shows eighty percent with one violation, which blocks release. I open that trajectory, see the agent misread the delivery date, fix the tool's date field description, and re-run until it's five out of five.

### Recap

To recap: evaluate agents on outcome, process and efficiency. Build realistic, adversarial task suites, run each task several times, and treat policy violations as failures. Read trajectories to find causes. And climb from prompt to autonomous agent only when evaluation proves you need to. Your next step is to take one agent idea from your organisation, place it on the ladder, and justify the lowest rung that would work. For more depth, continue with Building Production AI Agents and Evaluating and Monitoring LLM Applications.

### Try this now (20 minutes)

Try this now. Pick one agent idea from your organisation. Place it on the ladder: single prompt, prompt plus retrieval, fixed workflow, agent with approvals, or more autonomous agent. Write two sentences justifying the lowest rung that could work. Then write three test tasks for it, one normal, one with an obstacle and one adversarial, each with a success check and a forbidden action.

## Key takeaways

- Agents produce trajectories; evaluate outcomes, process and efficiency.
- Build sandboxed, realistic task suites with obstacles and adversarial cases; run each task multiple times.
- Read trajectories to find causes, often tool descriptions or missing instructions.
- Climb the ladder from single prompt to autonomous agent only when evaluation shows simpler designs fall short.

## Try it

Pick an 'agent' idea from your organisation. Place it on the ladder (prompt, RAG, workflow, agent with approvals, autonomous) and justify the lowest rung that would work.

- [Previous: Securing agents: the OWASP Top 10 for Agentic Applications](https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp/securing-agents-owasp-agentic)
- [All lessons of Latest AI Techniques: RAG, Tool Use, Agents & MCP](https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp)
