Building AI Products & WorkflowsEvaluation-driven development and production monitoring · Lesson 11 of 18

Evaluation-driven development

Article · 12 min · 8 min lecture

Video lecture

Evaluation-driven development

15 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 15

Evaluation-driven development

  • Why AI needs a different loop
  • Evaluation system components
  • Error analysis and staged rollouts

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why AI products need a different development loop

In traditional software, you specify behaviour and test it deterministically. AI features are probabilistic: the same input can produce different outputs, a prompt tweak can fix one case and break another, and a model update can shift behaviour overnight. Teams that succeed treat evaluation as the core of development, much as test-driven development treats tests.

The loop

1. Define success criteria (from the AI brief)
2. Build the evaluation set (real inputs + grading notes)
3. Build the simplest version that could work
4. Run evals -> read failures -> categorise
5. Change one thing (prompt, retrieval, model, tool, UX)
6. Re-run evals; keep changes that improve without regressions
7. Repeat until launch criteria are met; keep running after launch

Components of an evaluation system

  • Evaluation sets: representative, edge, failure and adversarial cases, tagged by category (see our Advanced Prompt Engineering course for depth).
  • Graders: code checks for objective criteria; rubric-based LLM judges, calibrated against humans, for subjective criteria; human review for samples.
  • A harness: a script or tool that runs the current version against the set and reports results by tag, with comparisons to the previous version.
  • Launch thresholds: agreed minimums for quality and safety metrics.
  • Regression gates: changes that drop key metrics below thresholds don't ship.

What product managers own

Evaluation is not only an engineering task. Product and domain experts should own:

  • Success criteria and rubrics: what "good" means for users.
  • Case selection: which scenarios matter most, including high-risk ones.
  • Grading calibration: expert judgements that LLM judges are compared with.
  • Trade-off decisions: is a small quality gain worth a latency increase?

Error analysis: the highest-leverage activity

Aggregate scores tell you how much; reading failures tells you why. Regularly review a sample of failures and categorise them:

Failure categoryCount (illustrative)Likely fix
Retrieved wrong document14Chunking, metadata filters
Ignored policy exception9Prompt instruction + examples
Wrong tone in Arabic replies7Examples, model choice for Arabic
Missing data in source system5Data fix, not AI fix

Fix the biggest category first. You'll often find some failures aren't AI problems at all (data quality, unclear policy, UX confusion).

Online evaluation: A/B tests and staged rollouts

Offline evaluation predicts; online evaluation confirms. Options:

  • Shadow mode: new version runs silently alongside; compare outputs.
  • Staged rollout: small percentage of traffic first, watching metrics and guardrails.
  • A/B tests: randomised comparison of versions on business outcomes (handling time, conversion, satisfaction).

For A/B tests, pre-define the primary metric, guardrails and duration, and avoid stopping early at the first favourable result.

Worked example

A support-reply drafting feature launches in pilot. Offline, it meets the quality threshold. Online, edit rates are higher than expected for returns questions. Error analysis shows agents rewrite drafts because they cite an outdated returns window. The fix is in retrieval (a superseded document) plus a regression case added to the evaluation set. Edit rates for returns fall the following week, and the same error can't silently return.

Common mistakes

  • Launching on "it looked good in the demo" with no evaluation set.
  • Evaluation sets built once and never refreshed.
  • Aggregate scores without error analysis.
  • Changing prompt, model and retrieval at the same time, so nobody knows what helped.

Hands-on: a minimal eval harness with a release gate

You do not need a platform to start. A folder of cases, a few graders and a comparison report will carry you a long way; move to a dedicated evaluation tool when volume and team size justify it.

evals/
  cases.jsonl          # {"id","input","tags":[...],"expect":{...}}
  graders.py           # code checks + LLM judge
  run.py               # runs a version, writes results/<version>.json, compares with baseline
# graders.py
import json, os, re
import anthropic

client = anthropic.Anthropic()
JUDGE_MODEL = os.environ.get("JUDGE_MODEL", "claude-opus-5")   # pin and record the judge version

def must_include(output, expect):          # objective, cheap, deterministic
    return all(s.lower() in output.lower() for s in expect.get("must_include", []))

def must_not_include(output, expect):
    return not any(re.search(p, output, re.I) for p in expect.get("must_not_match", []))

JUDGE_PROMPT = """Grade the RESPONSE against the RUBRIC. Think about each criterion,
then output only JSON: {{"pass": true|false, "failed_criteria": [..]}}.
RUBRIC: {rubric}
INPUT: {input}
RESPONSE: {output}"""

def judge(case, output):
    resp = client.messages.create(model=JUDGE_MODEL, max_tokens=800, messages=[{"role": "user",
        "content": JUDGE_PROMPT.format(rubric=case["expect"]["rubric"], input=case["input"], output=output)}])
    text = "".join(b.text for b in resp.content if b.type == "text")
    try:
        return json.loads(text[text.find("{"): text.rfind("}") + 1])
    except ValueError:
        return {"pass": False, "failed_criteria": ["judge_output_unparseable"]}
# run.py (core of the comparison)
import json, statistics, sys
from graders import must_include, must_not_include, judge

def evaluate(system_fn, version):
    cases = [json.loads(l) for l in open("evals/cases.jsonl", encoding="utf-8")]
    rows = []
    for c in cases:
        out = system_fn(c["input"])
        ok = must_include(out, c["expect"]) and must_not_include(out, c["expect"])
        if ok and "rubric" in c["expect"]:
            ok = judge(c, out)["pass"]
        rows.append({"id": c["id"], "tags": c["tags"], "pass": ok})
    json.dump(rows, open(f"results/{version}.json", "w"))
    return rows

def by_tag(rows):
    tags = {}
    for r in rows:
        for t in r["tags"]:
            tags.setdefault(t, []).append(r["pass"])
    return {t: statistics.mean(v) for t, v in tags.items()}

def gate(new_rows, base_rows, max_drop=0.03, floors={"safety": 1.0}):
    new, base = by_tag(new_rows), by_tag(base_rows)
    problems = [f"{t}: {base[t]:.2f} -> {new.get(t, 0):.2f}" for t in base if new.get(t, 0) < base[t] - max_drop]
    problems += [f"{t} below floor {f}" for t, f in floors.items() if new.get(t, 0) < f]
    print("\n".join(problems) or "PASS: no regressions")
    sys.exit(1 if problems else 0)

Calibrate the judge first: grade 50 outputs by hand, compare, and adjust the rubric until agreement is high. Record the judge model and prompt version with every result, because changing the judge changes the scores. The flagship course Evaluating and Monitoring LLM Applications goes deeper into datasets, judges, online evaluation and tooling.

Going further

Store every production interaction's inputs, outputs, versions and user signals (with privacy controls) in an evaluation store. Sample from it weekly to refresh your evaluation set, so it tracks how real usage evolves. The evaluation set is a living product asset.

Key takeaways

  • AI features are probabilistic; make evaluation the core development loop.
  • Components: evaluation sets, calibrated graders, a harness, launch thresholds and regression gates.
  • Product and domain experts own success criteria, case selection, grading calibration and trade-offs.
  • Error analysis finds the biggest failure categories; confirm offline results with shadow mode, staged rollouts and A/B tests.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. What is a regression gate?
  2. Error analysis shows many failures come from missing data in a source system. What is the right fix?
  3. Why should product experts own rubrics and case selection?

Put it into practice

Create an error-analysis table for an AI feature (or prototype) from 30 failures, categorised with likely fixes, and prioritise the top category.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.