Evaluating and Monitoring LLM ApplicationsFoundations: success criteria, error analysis and datasets · Lesson 2 of 16

Error analysis and failure taxonomies

Article · 13 min · 9 min lecture

Video lecture

Error analysis and failure taxonomies

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Error analysis and failure taxonomies

  • Read real outputs
  • Open and axial coding
  • From taxonomy to evals

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Look at your data first

The single highest-return activity in LLM evaluation is also the least glamorous: reading real outputs and writing down what went wrong. Practitioners such as Hamel Husain and Shreya Shankar, who teach evals widely, emphasize that teams who skip this step build metrics for problems they imagine rather than problems they have. Generic metrics ("helpfulness", "coherence") rarely match the specific ways your system fails.

Error analysis, step by step

This process borrows from qualitative research (open and axial coding):

  1. Sample traces. Pull 50–100 recent conversations or runs, including full context: user input, retrieved documents, tool calls, final output. Mix random samples with ones flagged by users.
  2. Open coding. For each trace, write a short free-text note on the first thing that went wrong, if anything. Be specific: "quoted UK return window to a UAE customer", not "wrong answer".
  3. Axial coding. Group notes into categories: a failure taxonomy. Merge duplicates, split vague buckets.
  4. Count. How often does each failure occur? Which are most severe?
  5. Prioritize. Frequency × severity × fixability. Some failures need a prompt change; some need retrieval fixes; some need a product decision.
  6. Build evals for the top categories, then fix, then re-measure.

Repeat after significant changes and periodically in production.

Worked example: a taxonomy from 80 traces

A property portal in Dubai runs an assistant that answers questions about listings and service charges. Two engineers reviewed 80 traces in an afternoon:

Failure categoryCountSeverityLikely fix
Uses listing data from wrong building (retrieval mix-up)11HighMetadata filter by building ID
Answers in English when user wrote Arabic7MediumLanguage instruction + eval
Invents service-charge figures when data missing5HighAbstain rule + groundedness eval
Over-long answers on mobile9LowLength constraint
Fails to hand off viewing requests to agent4HighTool-call check
No problem44––

Nobody had guessed that retrieval mix-ups between similarly named buildings were the top issue. A generic "answer quality" score would have hidden it.

Tooling for error analysis

You do not need special software to start: a spreadsheet with columns (trace link, input, output, note, category) works. As volume grows, observability platforms (Langfuse, LangSmith, Braintrust, Arize Phoenix and others) provide trace viewers, annotation queues and tagging. A lightweight custom viewer is often worth building: it should show the whole trace on one screen and make note-taking a keystroke.

Hands-on: export traces and create a coding sheet

If your traces are in JSON lines, a short script creates a review sheet:

# make_review_sheet.py: sample traces into a CSV for open coding
import csv, json, random

random.seed(7)
with open("traces.jsonl", encoding="utf-8") as f:
    traces = [json.loads(line) for line in f]

flagged = [t for t in traces if t.get("user_feedback") == "negative"]
others = [t for t in traces if t.get("user_feedback") != "negative"]
sample = flagged[:30] + random.sample(others, k=min(70, len(others)))

with open("review_sheet.csv", "w", newline="", encoding="utf-8") as f:
    w = csv.writer(f)
    w.writerow(["trace_id", "user_input", "retrieved_titles", "tool_calls", "output", "first_failure_note", "category"])
    for t in sample:
        w.writerow([
            t["id"], t["input"],
            "; ".join(d["title"] for d in t.get("retrieved", [])),
            "; ".join(c["name"] for c in t.get("tool_calls", [])),
            t["output"], "", "",
        ])
print(f"wrote {len(sample)} rows")

After coding, a pivot table gives you the taxonomy counts.

Can an LLM help with error analysis?

Yes, carefully. An LLM can propose categories from your notes, or pre-label traces with your taxonomy for you to confirm. It should not replace the initial reading: the insight comes from a human who knows the product seeing the failures. Always spot-check machine labels.

Who should do it

A domain expert (support lead, compliance officer, product owner) should be involved, ideally as the "benevolent dictator" who decides what counts as a failure. Engineers alone often miss domain errors; domain experts alone may miss systemic causes. Pair them.

Pitfalls

  • Coding every error in a trace. Note the first upstream failure; downstream errors are often consequences.
  • Vague categories ("bad answer"). If a category does not suggest a fix, split it.
  • Only reading complaints. Silent failures are common; include random samples.
  • Doing it once. Taxonomies drift as the product and users change.

How to measure success

A living failure taxonomy with counts, reviewed monthly, where every top category has a corresponding eval.

Key takeaways

  • Reading real outputs is the highest-return eval activity; generic metrics miss specific failures.
  • Sample 50–100 full traces, note the first failure specifically, group into a taxonomy, count and prioritize.
  • Involve a domain expert paired with an engineer; LLMs can assist but not replace the first reading.
  • Refresh the taxonomy regularly and build an eval for every top category.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. When open coding a trace with several errors, what should you note?
  2. Why include random samples, not just user-flagged traces?

Put it into practice

Sample 100 traces with the script in this lesson, open-code them, and produce a taxonomy table with counts and severity.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.