---
title: "Error analysis and failure taxonomies | Optimize All Academy"
description: "Look at your data first The single highest-return activity in LLM evaluation is also the least glamorous: reading real outputs and writing down what went…"
url: https://optimizeall.com/learn/llm-evals-and-observability/error-analysis-and-failure-taxonomies
updated: 2026-10-05
---

Evaluating and Monitoring LLM Applications · Foundations: success criteria, error analysis and datasets · lesson 2 of 16 · 13 min

# Error analysis and failure taxonomies

## Look at your data first

The single highest-return activity in LLM evaluation is also the least glamorous: **reading real outputs and writing down what went wrong**. Practitioners such as Hamel Husain and Shreya Shankar, who teach evals widely, emphasize that teams who skip this step build metrics for problems they imagine rather than problems they have. Generic metrics ("helpfulness", "coherence") rarely match the specific ways your system fails.

## Error analysis, step by step

This process borrows from qualitative research (open and axial coding):

1. **Sample traces.** Pull 50–100 recent conversations or runs, including full context: user input, retrieved documents, tool calls, final output. Mix random samples with ones flagged by users.
2. **Open coding.** For each trace, write a short free-text note on the *first* thing that went wrong, if anything. Be specific: "quoted UK return window to a UAE customer", not "wrong answer".
3. **Axial coding.** Group notes into categories: a **failure taxonomy**. Merge duplicates, split vague buckets.
4. **Count.** How often does each failure occur? Which are most severe?
5. **Prioritize.** Frequency × severity × fixability. Some failures need a prompt change; some need retrieval fixes; some need a product decision.
6. **Build evals** for the top categories, then fix, then re-measure.

Repeat after significant changes and periodically in production.

## Worked example: a taxonomy from 80 traces

A property portal in Dubai runs an assistant that answers questions about listings and service charges. Two engineers reviewed 80 traces in an afternoon:

| Failure category | Count | Severity | Likely fix |
|---|---|---|---|
| Uses listing data from wrong building (retrieval mix-up) | 11 | High | Metadata filter by building ID |
| Answers in English when user wrote Arabic | 7 | Medium | Language instruction + eval |
| Invents service-charge figures when data missing | 5 | High | Abstain rule + groundedness eval |
| Over-long answers on mobile | 9 | Low | Length constraint |
| Fails to hand off viewing requests to agent | 4 | High | Tool-call check |
| No problem | 44 | – | – |

Nobody had guessed that retrieval mix-ups between similarly named buildings were the top issue. A generic "answer quality" score would have hidden it.

## Tooling for error analysis

You do not need special software to start: a spreadsheet with columns (trace link, input, output, note, category) works. As volume grows, observability platforms (Langfuse, LangSmith, Braintrust, Arize Phoenix and others) provide trace viewers, annotation queues and tagging. A lightweight custom viewer is often worth building: it should show the whole trace on one screen and make note-taking a keystroke.

## Hands-on: export traces and create a coding sheet

If your traces are in JSON lines, a short script creates a review sheet:

```python
# make_review_sheet.py: sample traces into a CSV for open coding
import csv, json, random

random.seed(7)
with open("traces.jsonl", encoding="utf-8") as f:
    traces = [json.loads(line) for line in f]

flagged = [t for t in traces if t.get("user_feedback") == "negative"]
others = [t for t in traces if t.get("user_feedback") != "negative"]
sample = flagged[:30] + random.sample(others, k=min(70, len(others)))

with open("review_sheet.csv", "w", newline="", encoding="utf-8") as f:
    w = csv.writer(f)
    w.writerow(["trace_id", "user_input", "retrieved_titles", "tool_calls", "output", "first_failure_note", "category"])
    for t in sample:
        w.writerow([
            t["id"], t["input"],
            "; ".join(d["title"] for d in t.get("retrieved", [])),
            "; ".join(c["name"] for c in t.get("tool_calls", [])),
            t["output"], "", "",
        ])
print(f"wrote {len(sample)} rows")
```

After coding, a pivot table gives you the taxonomy counts.

## Can an LLM help with error analysis?

Yes, carefully. An LLM can propose categories from your notes, or pre-label traces with your taxonomy for you to confirm. It should not replace the initial reading: the insight comes from a human who knows the product seeing the failures. Always spot-check machine labels.

## Who should do it

A **domain expert** (support lead, compliance officer, product owner) should be involved, ideally as the "benevolent dictator" who decides what counts as a failure. Engineers alone often miss domain errors; domain experts alone may miss systemic causes. Pair them.

## Pitfalls

- **Coding every error in a trace.** Note the first upstream failure; downstream errors are often consequences.
- **Vague categories** ("bad answer"). If a category does not suggest a fix, split it.
- **Only reading complaints.** Silent failures are common; include random samples.
- **Doing it once.** Taxonomies drift as the product and users change.

## How to measure success

A living failure taxonomy with counts, reviewed monthly, where every top category has a corresponding eval.

## Video lecture: Error analysis and failure taxonomies

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. Error analysis and failure taxonomies
2. Analogy: reading the actual charts
3. Why start here
4. Steps 1–2
5. Steps 3–6
6. Case: Dubai property assistant
7. Tooling
8. LLM assistance + people
9. Pitfalls
10. Example: three open codes
11. Scenario: 100 pharmacy traces (illustrative)
12. Common mistakes
13. Deeper: sharpening a vague note
14. Watch me do it: coding 100 traces
15. Recap

## Lecture transcript

### Error analysis and failure taxonomies

If you only do one thing from this entire course, make it this: read your system's outputs and write down what went wrong. It sounds too simple to matter. It is the highest-return activity in LLM evaluation, and most teams skip it. In this lecture you will learn a practical error-analysis process, how to turn notes into a failure taxonomy, and how that taxonomy tells you exactly which evals to build.

### Analogy: reading the actual charts

Here is an analogy for error analysis. A doctor who only ever looks at your average blood pressure over a year would miss the dangerous spike last Tuesday. Good doctors look at the actual readings, notice patterns, and name what they see before choosing a treatment. Error analysis is looking at the actual readings of your AI system, one trace at a time, and naming what went wrong, before you decide which tests and fixes to build.

### Why start here

Why does this matter so much? Because generic metrics like helpfulness or coherence rarely match the specific ways your system fails. Practitioners who teach evals widely, such as Hamel Husain and Shreya Shankar, stress the same point. Teams that skip error analysis build metrics for problems they imagine, not problems they have. And then they optimize a number that does not move the user experience.

### Steps 1–2

The process has six steps. First, sample fifty to a hundred recent traces with full context: input, retrieved documents, tool calls and output. Mix random samples with ones users flagged. Second, open coding: for each trace, write a short, specific note about the first thing that went wrong. Not wrong answer, but quoted the UK return window to a UAE customer.

### Steps 3–6

Third, axial coding: group your notes into categories, merging duplicates and splitting vague buckets. That becomes your failure taxonomy. Fourth, count how often each category occurs and how severe it is. Fifth, prioritize by frequency, severity and how fixable it is. Sixth, build evals for the top categories, fix, and re-measure. Then repeat after big changes, and regularly in production.

### Case: Dubai property assistant

Here is what it looks like in practice. A property portal in Dubai reviewed eighty traces from its listings assistant in one afternoon. The top failure was a surprise: answers mixing up data from similarly named buildings, a retrieval problem. Others included replying in English to Arabic questions, inventing service-charge figures when data was missing, and failing to hand off viewing requests. Nobody had guessed the building mix-up. A generic quality score would have hidden it completely.

### Tooling

You do not need special tools to start. A spreadsheet with columns for trace link, input, output, note and category works fine, and the lesson includes a script that samples traces into exactly that sheet. As volume grows, observability platforms offer trace viewers and annotation queues. Many teams also build a tiny custom viewer that shows a whole trace on one screen and makes adding a note a single keystroke. Friction is the enemy of looking at data.

### LLM assistance + people

Can a language model help? Yes, carefully. It can suggest categories from your notes, or pre-label new traces with your existing taxonomy for you to confirm. But it should not replace the first reading. The insight comes from someone who knows the product seeing the failures with their own eyes. And that someone should include a domain expert, like a support lead or compliance officer, paired with an engineer. Each catches what the other misses.

### Pitfalls

Four pitfalls. Coding every error in a trace instead of the first upstream failure, because downstream errors are often just consequences. Vague categories: if a category does not suggest a fix, split it. Only reading complaints: silent failures are common, so always include random samples. And doing it once: your taxonomy drifts as your product and users change.

### Example: three open codes

A simple example of open coding. Trace one: a customer in Karachi asks about cash on delivery, and the bot explains card payments. Note: answered a different payment method than asked. Trace two: a customer asks in Urdu, and the bot replies in English. Note: wrong response language. Trace three: the bot quotes a delivery fee that is not in the policy. Note: invented fee. Three traces, three specific notes. After eighty traces, you group them, and you might find invented fees appear far more often than anyone suspected.

### Scenario: 100 pharmacy traces (illustrative)

Now a realistic scenario with illustrative numbers. An online pharmacy in Riyadh reviews one hundred traces from its product-question assistant. Sixty-one have no problem. Of the rest, fourteen recommend a product that is out of stock, nine answer in English to Arabic questions, eight give dosage-adjacent advice the policy forbids, and eight are other issues. The dosage cases are fewer than the stock cases, but far more severe, so they become the first blocking eval. The stock issue turns out to be a data freshness problem, fixed without touching the prompt at all.

### Common mistakes

Common mistakes. Writing notes that are too vague to act on, like bad answer. Letting the first reviewer's categories become permanent without revisiting them. Skipping the domain expert because the engineers feel confident. And treating error analysis as a one-time project instead of a monthly habit. Here is a question for you: when did someone on your team last read fifty real outputs from your AI feature, start to finish?

### Deeper: sharpening a vague note

One level deeper on a vague note. Suppose a reviewer writes: answer was wrong. Ask two follow-up questions. Wrong compared with what source? And what was the first step that went wrong, retrieval, reasoning or formatting? The note becomes: quoted the old return window because the retrieved policy was last year's version. Now it points at a fix, the index refresh.

### Watch me do it: coding 100 traces

Watch me do it. I run the review-sheet script from the lesson on last week's traces. It takes thirty user-flagged conversations and seventy random ones, and writes a CSV with the input, the titles of retrieved documents, tool calls and output, plus two empty columns: first failure note, and category. I open it and start reading. Row one: fine, I leave the note empty. Row two: a customer in Dubai asked about returns and the retrieved document titles show the UK policy. My note: retrieved UK returns policy for a UAE customer. Row three: fine. Row four: the customer wrote in Urdu, the answer is in English. Note: replied in English to an Urdu message. Row five: the bot quoted a delivery fee, and no retrieved document mentions fees. Note: invented delivery fee, not in context. I keep going for about ninety minutes. Then I sort the notes and group them. Wrong-country policy retrieved: twelve. Wrong response language: seven. Invented fees or dates: six. Missed escalation: three. Now the categories suggest their fixes: a country filter in retrieval, a language instruction plus a check, an abstain rule plus a groundedness eval, and a code check on the escalation tool call. The invented-fee category goes first because it is the most severe.

### Recap

Recap. Read your data before choosing metrics. Sample, open code, group into a taxonomy, count, prioritize and build evals for the top categories. Involve a domain expert. Your next step: use the script in the lesson to sample one hundred traces from your system, code them this week, and bring the taxonomy to your next planning meeting.

## Key takeaways

- Reading real outputs is the highest-return eval activity; generic metrics miss specific failures.
- Sample 50–100 full traces, note the first failure specifically, group into a taxonomy, count and prioritize.
- Involve a domain expert paired with an engineer; LLMs can assist but not replace the first reading.
- Refresh the taxonomy regularly and build an eval for every top category.

## Try it

Sample 100 traces with the script in this lesson, open-code them, and produce a taxonomy table with counts and severity.

- [Previous: Why evals, and defining success](https://optimizeall.com/learn/llm-evals-and-observability/why-evals-and-defining-success)
- [Next: Building eval datasets: golden, production and synthetic](https://optimizeall.com/learn/llm-evals-and-observability/building-eval-datasets)
- [All lessons of Evaluating and Monitoring LLM Applications](https://optimizeall.com/learn/llm-evals-and-observability)
