Skip to content

Evaluating and Monitoring LLM Applications · Foundations: success criteria, error analysis and datasets · lesson 1 of 16 · 13 min

Why evals, and defining success

Vibes do not scale

Most LLM applications begin with "vibe checks": a developer tries ten prompts, the answers look good, and the feature ships. Then a prompt tweak fixes one complaint and silently breaks three other cases, a model upgrade changes tone, or a new document set confuses retrieval. Without evals, every change is a gamble and every debate about quality is an opinion contest.

An eval is a repeatable measurement of how well your LLM system performs on a defined task, using a dataset of inputs and a way to score outputs. Evals do for LLM apps what unit and integration tests do for conventional software, with one big difference: outputs are non-deterministic and often have many acceptable answers, so scoring is itself a design problem.

Why evals are the core engineering asset

  • Safe iteration. Change prompts, models, retrieval or tools and know within minutes what improved and what regressed.
  • Model choice and migration. Compare vendors and model versions on your task, not public leaderboards; upgrade with confidence when a model is deprecated.
  • Cost optimization. Prove a cheaper or smaller model is good enough for a route.
  • Shared language. Product, engineering, legal and support agree on what "good" means.
  • Governance. Documented evaluation supports risk management frameworks (for example NIST's AI RMF "Measure" function) and customer due diligence.

Defining success before measuring it

Start with the product, not the metric. For each feature, write down:

  1. The job. What the user is trying to accomplish.
  2. Quality dimensions. Usually several: correctness, groundedness in sources, completeness, tone and brand voice, safety, format validity, latency, cost.
  3. Must-never-happen failures. Hard constraints: leaking another customer's data, promising refunds policy does not allow, giving medical dosage advice, inventing prices.
  4. Thresholds. What score on each dimension is good enough to ship, and which are blocking.

Worked example: success criteria for a support bot

An e-commerce brand selling across Pakistan, the UAE and the UK is launching a support assistant.

| Dimension | Definition | Measured by | Ship threshold | |---|---|---|---| | Resolution correctness | Answer matches policy for the customer's country | LLM judge with rubric, calibrated on 100 human-labeled cases | ≥ 90% | | Groundedness | Every factual claim supported by retrieved policy text | Groundedness judge | ≥ 95% | | Escalation | Hands off to a human for refunds over limit, legal threats, or abuse | Code check on tool call | 100% on test set | | Privacy | Never reveals order details without verification | Adversarial test set, code + judge | 0 failures (blocking) | | Tone | Friendly, concise, matches brand guide | Pairwise judge vs reference | Not worse than current | | Latency | Time to first token | Tracing | p95 ≤ agreed target | | Cost | Tokens per conversation | Tracing | ≤ budget |

Note the mix: code-based checks where possible, LLM judges where needed, human labels to calibrate the judges, and production telemetry for latency and cost. The numbers are illustrative; set yours with stakeholders.

Eval-driven development

The workflow mirrors test-driven development:

  1. Collect real or realistic examples, and look at the outputs (next lesson: error analysis).
  2. Write evals for the failure modes that matter.
  3. Change the system (prompt, retrieval, model, tools).
  4. Run evals; compare with the baseline; inspect regressions.
  5. Ship when thresholds are met; monitor in production; feed new failures back into the dataset.

Evals are never "done". Production reveals failure modes you did not imagine, and each becomes a new test case.

Hands-on: an eval spec template

# Eval spec: <feature>
## Job to be done
## Quality dimensions (definition, how measured, threshold, blocking?)
| Dimension | Definition | Method | Threshold | Blocking |
## Must-never-happen failures (each needs a dedicated test set)
## Datasets (golden, synthetic, production sample; size; owner; refresh cadence)
## Baseline (current system scores, date, model/prompt version)
## Decision rule (e.g. ship if all blocking pass and no dimension regresses > 2 points)

Pitfalls

  • One overall score. Averages hide the failure that matters; track dimensions separately.
  • Metric before meaning. Choosing BLEU or cosine similarity because a library offers it, not because it reflects quality.
  • Leaderboard thinking. Public benchmarks rarely predict performance on your task.
  • No owner. Evals rot without someone responsible for datasets and thresholds.

How to measure success

You have succeeded when every prompt or model change runs against the eval suite before release, and quality debates are settled by looking at the results together.

Video lecture: Why evals, and defining success

Lecture coming soon · 14 chapters · about 8 minutes. Read the full transcript below.

  1. Why evals, and defining success
  2. Analogy: the tasting spoon
  3. What is an eval?
  4. Why evals matter
  5. Define success first
  6. Example: support bot
  7. Mix your methods
  8. Eval-driven development
  9. Pitfalls
  10. Example: product-title bot
  11. Scenario: cheaper model? (illustrative)
  12. Deeper: thresholds mirror error cost
  13. Watch me do it: an eval spec
  14. Recap

Lecture transcript

Why evals, and defining success

Here is a story nearly every AI team knows. A developer tweaks the prompt to fix one customer complaint. It works. Two weeks later, support notices the bot has started inventing delivery dates for UK orders. Nobody connected the two, because nobody was measuring. In this lecture you will learn what an eval is, why it is the core engineering asset of any LLM product, and how to define success before you measure it.

Analogy: the tasting spoon

Here is an analogy. Imagine running a restaurant where the chef changes recipes every week but nobody ever tastes the food before it goes out. Some weeks the new recipe is better. Some weeks it is quietly worse, and you only find out when regulars stop coming. An eval suite is the tasting spoon. It does not cook for you. It just makes sure nothing leaves the kitchen without someone checking that it still tastes the way it should.

What is an eval?

An eval is a repeatable measurement of how well your system performs a defined task. It has three parts: a dataset of inputs, the system under test, and a way to score the outputs. It is the LLM equivalent of a test suite, with one crucial twist. Outputs are non-deterministic, and there are often many acceptable answers. So how you score is itself a design problem, and much of this course is about solving it well.

Why evals matter

Why are evals the core asset? First, safe iteration: change a prompt, model or retriever and know in minutes what improved and what regressed. Second, model choice: compare vendors and versions on your task, not public leaderboards, and migrate confidently when a model is retired. Third, cost: prove a cheaper model is good enough for a given route. Fourth, a shared language for product, engineering and legal. And fifth, governance: documented measurement supports frameworks like the NIST AI Risk Management Framework and customer due diligence.

Define success first

Now, defining success. Start with the product, not the metric. What is the user's job? Which quality dimensions matter: correctness, groundedness in sources, completeness, tone, safety, format, latency and cost? Which failures must never happen, like leaking another customer's data or promising a refund that policy does not allow? And what threshold on each dimension is good enough to ship?

Example: support bot

Let's make it concrete with a support bot for an e-commerce brand selling across Pakistan, the UAE and the UK. Resolution correctness means the answer matches policy for the customer's country, scored by a calibrated LLM judge. Groundedness means every claim is supported by retrieved policy text. Escalation means refunds over the limit or legal threats go to a human, checked in code by looking at the tool call. Privacy failures are blocking, with zero tolerated. Then tone, latency and cost.

Mix your methods

Notice the mix of methods. Use code-based checks wherever you can, because they are cheap and deterministic. Use LLM judges where judgment is needed. Use human labels to calibrate those judges. And use production telemetry for latency and cost. No single method covers everything, and pretending one score does is the fastest way to miss the failure that matters.

Eval-driven development

The workflow is eval-driven development. Collect examples and actually look at the outputs. Write evals for the failure modes that matter. Change the system. Run the evals and compare against the baseline, looking closely at regressions, not just the average. Ship when thresholds are met. Then monitor production and feed every new failure back into the dataset. Evals are never done, because users will always find cases you did not imagine.

Pitfalls

A few pitfalls. One overall score hides the failure that matters, so track dimensions separately. Choosing a metric because a library offers it, rather than because it reflects quality. Trusting public leaderboards to predict your task. And having no owner, because eval datasets rot without someone responsible for them.

Example: product-title bot

A simple example of defining success. Feature: a bot that writes product titles for an online shop. The job: short, searchable titles. Dimensions: includes brand and product type, under seventy characters, no invented features, correct language. Must never happen: claiming a certification the product does not have. Thresholds: ninety-five percent of titles pass the length and brand checks in code, and zero invented certifications on a test set of sixty products with known specs. That took five minutes, and now every prompt change can be judged against it.

Scenario: cheaper model? (illustrative)

Now a realistic scenario with illustrative numbers. A travel agency in Lahore wants to switch its itinerary assistant to a cheaper model. Without evals, the debate goes on for weeks. With an eval spec, they run both models on one hundred and twenty real requests. The cheaper model matches on formatting and tone, costs much less, but gets visa requirements wrong noticeably more often, a blocking dimension. Decision in one afternoon: use the cheaper model for itinerary formatting, keep the stronger model for visa questions. Evals turned an opinion fight into a routing decision.

Deeper: thresholds mirror error cost

One level deeper on thresholds. Why ninety percent for correctness but zero failures for privacy? Because the cost of each error differs. A slightly wrong answer about delivery times costs a follow-up message. A privacy leak costs trust, and in many jurisdictions, legal obligations. Thresholds should mirror the cost of being wrong, and you should write that reasoning next to each number.

Watch me do it: an eval spec

Watch me do it. I'm filling in the eval spec template for the multi-country support bot. Job to be done: customers in Pakistan, the UAE and the UK get correct answers about returns, delivery and orders without waiting for an agent. Quality dimensions, one row each. Resolution correctness: the answer matches policy for the customer's country; measured by a policy judge calibrated on one hundred human labels; threshold ninety percent; not blocking. Groundedness: every factual claim supported by retrieved policy text; groundedness judge; ninety-five percent. Escalation: refunds over the limit, legal threats and abuse go to a human; checked in code by looking for the handoff tool call; one hundred percent on the test set; blocking. Privacy: never reveals order details before verification; adversarial set with code and judge checks; zero failures; blocking. Tone: pairwise against current answers; not worse. Latency and cost: from tracing. Next, must-never-happen failures, each with its own test set: privacy leaks, and promising refunds outside policy. Datasets: sixty golden items from the support lead, one hundred production samples each month, and forty synthetic edge cases. Baseline: current scores with today's date and model version. And the decision rule, the line everyone signs: ship only if all blocking checks pass and no dimension drops by more than two points.

Recap

Recap. An eval is a dataset, a system and a scorer. It is your core asset for iteration, model choice, cost and governance. Define success from the product down, with dimensions, must-never-happen failures and thresholds. Your next step: fill in the eval spec template from the lesson for one feature you own, and agree the blocking criteria with your stakeholders.

Key takeaways

  • An eval is a dataset, a system under test and a scoring method; outputs are non-deterministic, so scoring is a design problem.
  • Evals enable safe iteration, model choice, cost optimization, shared language and governance evidence.
  • Define success from the product: job, quality dimensions, must-never-happen failures and thresholds.
  • Mix code checks, calibrated LLM judges, human labels and telemetry; never rely on one averaged score.

Try it

Complete the eval spec template for one LLM feature you own and agree the blocking criteria with a stakeholder.