Building AI Products & WorkflowsEvaluation-driven development and production monitoring · Lesson 11 of 18
Evaluation-driven development
Video lecture
Evaluation-driven development
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Evaluation-driven development
In normal software, you write a test once and it passes or fails the same way forever. AI features aren't like that. The same input can produce different outputs. A prompt tweak can fix one case and break three others. A provider model update can shift behaviour overnight. The teams that ship reliable AI treat evaluation as the engine of development. In this lesson you'll learn the loop, the parts of an evaluation system, what product managers own, and how error analysis tells you exactly what to fix.
0:38 Analogy: the tasting spoon
Here's an analogy. Building AI without evaluations is like cooking for a restaurant without tasting. You might follow the recipe perfectly, but ingredients vary, ovens differ, and one day the supplier changes the flour. Tasting every batch is how chefs stay consistent. Your evaluation set is the tasting spoon, and you use it every time anything changes.
1:03 The loop
Here's the loop. Define success criteria from your AI brief. Build an evaluation set of real inputs with grading notes. Build the simplest version that could work. Run the evaluation, read the failures and categorise them. Change one thing: the prompt, retrieval, model, a tool or the UX. Re-run, and keep changes that improve results without regressions. Repeat until you meet launch criteria, and keep running after launch. It's test-driven development, adapted for probabilistic systems.
1:36 Five components
An evaluation system has five parts. Evaluation sets with representative, edge, failure and adversarial cases, tagged by category. Graders: code checks for objective criteria, rubric-based LLM judges calibrated against humans for subjective ones, and human review for samples. A harness that runs a version against the set and reports results by tag, compared with the previous version. Launch thresholds agreed in advance. And regression gates, so changes that drop key metrics below thresholds don't ship.
2:09 What product owns
Evaluation isn't only an engineering job. Product managers and domain experts should own what good means for users, which scenarios matter most, including high-risk ones, the expert grades that LLM judges are calibrated against, and trade-off calls, like whether a small quality gain is worth extra latency. If the people who understand the customer don't own the rubric, you'll optimise the wrong thing very efficiently.
2:37 Error analysis → online checks
Aggregate scores tell you how much. Reading failures tells you why. Take a sample of failures and categorise them: retrieved the wrong document, ignored a policy exception, wrong tone in Arabic replies, missing data in the source system. Fix the biggest category first. You'll often discover some failures aren't AI problems at all. They're data quality, unclear policy or confusing UX. Then confirm offline gains online: shadow mode, where the new version runs silently alongside, staged rollouts to a small share of traffic, and A/B tests with a primary metric, guardrails and duration decided before you start.
3:19 Simple example: a friendlier prompt
A simple example. You change one line in a prompt to make replies friendlier. Your evaluation set of sixty cases shows tone scores rise, but three refund-policy answers now leave out the thirty-day window. Without the set, you'd have shipped a friendlier assistant that gets policy wrong. With it, you add an instruction about policy details, re-run, and ship a version that's friendlier and still accurate.
3:48 Worked example: support drafts
Here's a worked example. A support-reply drafting feature meets its offline quality threshold and launches in pilot. Online, agents edit returns-related drafts far more than expected. Error analysis shows the drafts cite an outdated returns window, because a superseded document is still indexed. The fix is in retrieval, not the prompt, plus a regression case added to the evaluation set. Edit rates for returns fall the next week, and that error can never silently return.
4:21 Business example (illustrative)
Illustrative numbers for the support drafts. Offline, the pilot passed eighty-eight percent of a hundred and fifty cases. Online, returns drafts were edited in about sixty percent of cases versus twenty-five percent elsewhere. After removing the superseded document and adding three regression cases, returns edits fell to about thirty percent in a week, and the regression cases have passed on every release since.
4:48 Hands-on in the lesson
The hands-on section gives you a minimal harness you can run without buying a platform: cases in a JSON lines file, cheap code graders for must-include and must-not-match rules, an LLM judge that returns pass or fail with failed criteria, results saved per version, and a release gate that fails the build if any tag drops more than a few points or safety falls below its floor. Calibrate the judge against fifty human-graded outputs first, and record the judge's model and prompt version with every result.
5:25 Common mistakes
Common mistakes. Building the evaluation set after launch. Using only easy, typical cases. Trusting an LLM judge without checking it against human grades. Changing several things at once. Letting the evaluation set go stale while real usage evolves. And keeping evaluation in an engineer's laptop, so product and domain experts never see the results or influence the rubric.
5:50 How you'll know it's working
How will you know evaluation-driven development is working? Every change has a before-and-after result on record. Regressions are caught by the gate, not by customers. The share of failures that surprise you in production shrinks over time, because each one becomes a new test case. And product conversations shift from I think it's better to here's what changed, by tag.
6:16 Watch me do it: harness + gate
Watch me do it. I open the harness folder. Cases dot JSON L has one case per line with an ID, input, tags and expectations. In graders, must include checks required phrases, must not include checks forbidden patterns, and judge sends the rubric, input and output to a pinned judge model and parses pass or fail, treating unparseable output as a fail. In run dot py, evaluate runs my system on every case, applies code checks first, calls the judge only if a rubric exists, and saves results per version. By tag computes pass rates. Gate compares the new version with the baseline and fails if any tag drops more than three points or safety falls below one hundred percent. I run version eight against version seven. The output says returns, point nine one to point eight four, and exits with a failure, so the change doesn't ship until I fix it.
7:22 Recap
To recap: AI features are probabilistic, so evaluation is the core development loop. Build tagged sets, calibrated graders, a harness, thresholds and gates. Product owns what good means. Error analysis finds the biggest failure categories, and staged rollouts confirm offline gains. Your next step is to categorise thirty failures from a feature or prototype, estimate the fix for each category, and start with the biggest. For more depth, see Evaluating and Monitoring LLM Applications.
7:54 Try this now (45 minutes)
Try this now. Create a folder with a cases file and write fifteen cases for one AI feature: ten typical, three edge cases and two that should be refused or escalated. For each, write what must be included and what must not appear. Run your current version, record pass or fail, and categorise every failure. That's your baseline and your first improvement backlog.
Why AI products need a different development loop
In traditional software, you specify behaviour and test it deterministically. AI features are probabilistic: the same input can produce different outputs, a prompt tweak can fix one case and break another, and a model update can shift behaviour overnight. Teams that succeed treat evaluation as the core of development, much as test-driven development treats tests.
The loop
1. Define success criteria (from the AI brief)
2. Build the evaluation set (real inputs + grading notes)
3. Build the simplest version that could work
4. Run evals -> read failures -> categorise
5. Change one thing (prompt, retrieval, model, tool, UX)
6. Re-run evals; keep changes that improve without regressions
7. Repeat until launch criteria are met; keep running after launchComponents of an evaluation system
- Evaluation sets: representative, edge, failure and adversarial cases, tagged by category (see our Advanced Prompt Engineering course for depth).
- Graders: code checks for objective criteria; rubric-based LLM judges, calibrated against humans, for subjective criteria; human review for samples.
- A harness: a script or tool that runs the current version against the set and reports results by tag, with comparisons to the previous version.
- Launch thresholds: agreed minimums for quality and safety metrics.
- Regression gates: changes that drop key metrics below thresholds don't ship.
What product managers own
Evaluation is not only an engineering task. Product and domain experts should own:
- Success criteria and rubrics: what "good" means for users.
- Case selection: which scenarios matter most, including high-risk ones.
- Grading calibration: expert judgements that LLM judges are compared with.
- Trade-off decisions: is a small quality gain worth a latency increase?
Error analysis: the highest-leverage activity
Aggregate scores tell you how much; reading failures tells you why. Regularly review a sample of failures and categorise them:
| Failure category | Count (illustrative) | Likely fix |
|---|---|---|
| Retrieved wrong document | 14 | Chunking, metadata filters |
| Ignored policy exception | 9 | Prompt instruction + examples |
| Wrong tone in Arabic replies | 7 | Examples, model choice for Arabic |
| Missing data in source system | 5 | Data fix, not AI fix |
Fix the biggest category first. You'll often find some failures aren't AI problems at all (data quality, unclear policy, UX confusion).
Online evaluation: A/B tests and staged rollouts
Offline evaluation predicts; online evaluation confirms. Options:
- Shadow mode: new version runs silently alongside; compare outputs.
- Staged rollout: small percentage of traffic first, watching metrics and guardrails.
- A/B tests: randomised comparison of versions on business outcomes (handling time, conversion, satisfaction).
For A/B tests, pre-define the primary metric, guardrails and duration, and avoid stopping early at the first favourable result.
Worked example
A support-reply drafting feature launches in pilot. Offline, it meets the quality threshold. Online, edit rates are higher than expected for returns questions. Error analysis shows agents rewrite drafts because they cite an outdated returns window. The fix is in retrieval (a superseded document) plus a regression case added to the evaluation set. Edit rates for returns fall the following week, and the same error can't silently return.
Common mistakes
- Launching on "it looked good in the demo" with no evaluation set.
- Evaluation sets built once and never refreshed.
- Aggregate scores without error analysis.
- Changing prompt, model and retrieval at the same time, so nobody knows what helped.
Hands-on: a minimal eval harness with a release gate
You do not need a platform to start. A folder of cases, a few graders and a comparison report will carry you a long way; move to a dedicated evaluation tool when volume and team size justify it.
evals/
cases.jsonl # {"id","input","tags":[...],"expect":{...}}
graders.py # code checks + LLM judge
run.py # runs a version, writes results/<version>.json, compares with baseline# graders.py
import json, os, re
import anthropic
client = anthropic.Anthropic()
JUDGE_MODEL = os.environ.get("JUDGE_MODEL", "claude-opus-5") # pin and record the judge version
def must_include(output, expect): # objective, cheap, deterministic
return all(s.lower() in output.lower() for s in expect.get("must_include", []))
def must_not_include(output, expect):
return not any(re.search(p, output, re.I) for p in expect.get("must_not_match", []))
JUDGE_PROMPT = """Grade the RESPONSE against the RUBRIC. Think about each criterion,
then output only JSON: {{"pass": true|false, "failed_criteria": [..]}}.
RUBRIC: {rubric}
INPUT: {input}
RESPONSE: {output}"""
def judge(case, output):
resp = client.messages.create(model=JUDGE_MODEL, max_tokens=800, messages=[{"role": "user",
"content": JUDGE_PROMPT.format(rubric=case["expect"]["rubric"], input=case["input"], output=output)}])
text = "".join(b.text for b in resp.content if b.type == "text")
try:
return json.loads(text[text.find("{"): text.rfind("}") + 1])
except ValueError:
return {"pass": False, "failed_criteria": ["judge_output_unparseable"]}# run.py (core of the comparison)
import json, statistics, sys
from graders import must_include, must_not_include, judge
def evaluate(system_fn, version):
cases = [json.loads(l) for l in open("evals/cases.jsonl", encoding="utf-8")]
rows = []
for c in cases:
out = system_fn(c["input"])
ok = must_include(out, c["expect"]) and must_not_include(out, c["expect"])
if ok and "rubric" in c["expect"]:
ok = judge(c, out)["pass"]
rows.append({"id": c["id"], "tags": c["tags"], "pass": ok})
json.dump(rows, open(f"results/{version}.json", "w"))
return rows
def by_tag(rows):
tags = {}
for r in rows:
for t in r["tags"]:
tags.setdefault(t, []).append(r["pass"])
return {t: statistics.mean(v) for t, v in tags.items()}
def gate(new_rows, base_rows, max_drop=0.03, floors={"safety": 1.0}):
new, base = by_tag(new_rows), by_tag(base_rows)
problems = [f"{t}: {base[t]:.2f} -> {new.get(t, 0):.2f}" for t in base if new.get(t, 0) < base[t] - max_drop]
problems += [f"{t} below floor {f}" for t, f in floors.items() if new.get(t, 0) < f]
print("\n".join(problems) or "PASS: no regressions")
sys.exit(1 if problems else 0)Calibrate the judge first: grade 50 outputs by hand, compare, and adjust the rubric until agreement is high. Record the judge model and prompt version with every result, because changing the judge changes the scores. The flagship course Evaluating and Monitoring LLM Applications goes deeper into datasets, judges, online evaluation and tooling.
Going further
Store every production interaction's inputs, outputs, versions and user signals (with privacy controls) in an evaluation store. Sample from it weekly to refresh your evaluation set, so it tracks how real usage evolves. The evaluation set is a living product asset.
Key takeaways
- AI features are probabilistic; make evaluation the core development loop.
- Components: evaluation sets, calibrated graders, a harness, launch thresholds and regression gates.
- Product and domain experts own success criteria, case selection, grading calibration and trade-offs.
- Error analysis finds the biggest failure categories; confirm offline results with shadow mode, staged rollouts and A/B tests.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Create an error-analysis table for an AI feature (or prototype) from 30 failures, categorised with likely fixes, and prioritise the top category.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.