AI Product Management: From Idea to Reliable AI Features · Evals, launch and impact measurement · lesson 11 of 16 · 16 min
Evals-driven development with your engineering team
The core loop of AI product development
In AI products, evaluations (evals) are the specification, the test suite and the progress report at once. Evals-driven development means every change (prompt, model, retrieval, tool, UI flow) is judged against the same evaluation set before it ships. It replaces "the demo looked good" with "the golden set score went from 3.9 to 4.3 with no regressions".
The loop
1. Define: golden set + rubric + thresholds (from the PRD)
2. Baseline: score the current system
3. Analyse errors: read failures, group them into categories
4. Hypothesise: pick the biggest failure category and a fix
5. Change: prompt, retrieval, model, tool, data, UX
6. Re-run evals: compare with baseline, check slices and regressions
7. Ship or iterate; add new real-world failures to the golden set
Error analysis: the step teams skip
Scores tell you how much is wrong; reading failures tells you what is wrong. For 50 failed examples, tag each with a failure category:
| Category | Count | Example | Likely fix | |---|---|---|---| | Retrieval missed the right document | 18 | Question about "EOS" (end-of-service) benefits missed the gratuity policy | Add synonyms/metadata; hybrid search | | Correct docs, wrong reasoning | 9 | Mixed up probation and notice periods | Prompt with explicit steps; examples | | Outdated source | 8 | Cited 2023 travel policy | Remove old versions; date filters | | Tone/format | 7 | Too long for WhatsApp | Output spec; length limit | | Should have abstained | 5 | Answered a question not covered | Stronger abstention instruction; threshold | | Language issue | 3 | Arabic answer to English question | Language detection rule |
Now the roadmap is obvious: fix retrieval first. PMs are often the best people to run error analysis because they understand users and domain.
Types of evals and who owns what
- Code-based checks (fast, cheap): JSON validity, required fields, length, forbidden phrases, citation present. Engineers own.
- Reference-based checks: compare with expected answers (exact match for labels, key facts for answers). Shared.
- Rubric scoring by humans: domain experts score a sample against the rubric. PM organises, experts score.
- LLM-as-judge: a model scores outputs with the rubric, calibrated against human scores (check agreement on 50+ items before trusting it; swap order in pairwise comparisons).
- Online signals: user feedback, edits, escalations, task completion. PM and analytics own.
Evals in the development workflow
- Run fast code checks and a small eval subset on every change (in CI).
- Run the full golden set before every release and on any model or provider change.
- Keep a dashboard: score over time by slice, cost per task, latency.
- Define regression gates: a release is blocked if a slice drops beyond the agreed tolerance.
Several open-source and commercial tools support this (for example promptfoo, Inspect, and LLM observability platforms with evaluation features); the choice matters less than the discipline.
Hands-on: an eval spec PMs and engineers can share
# evals/support-assistant/spec.yaml
golden_set: evals/support-assistant/golden_v1.3.jsonl # 240 cases, tagged by language, topic, difficulty
checks:
- name: json_valid
type: code
threshold: 1.0
- name: has_citation_for_claims
type: code
threshold: 0.98
- name: rubric_score
type: llm_judge
judge_prompt: evals/support-assistant/judge_rubric.md
calibration: 60 human-scored items, agreement >= 0.8 within 1 point
threshold: {mean: 4.2, max_share_score_1: 0.03}
- name: abstention_on_unanswerable
type: reference
subset: unanswerable
threshold: 0.9
slices: [language, topic, channel]
regression_gate: "no slice mean drops by more than 0.2 vs production"
run: {on_pull_request: smoke_subset_40, before_release: full, on_model_change: full}
owners: {golden_set: pm, checks: eng-lead, human_scoring: support-qa}
Worked example: a KSA HR assistant, week by week (illustrative)
- Week 1: baseline rubric mean 3.6. Error analysis: 40% of failures are retrieval misses on Arabic terms and abbreviations.
- Week 2: added Arabic synonyms and hybrid search; mean 3.9, Arabic slice +0.5, no regressions.
- Week 3: removed outdated policy versions; mean 4.1.
- Week 4: tightened abstention; unanswerable accuracy 72% → 91%; mean 4.2; launch gate met.
- Post-launch: 30 real failures from feedback added to golden set v1.4.
Pitfalls
- Judging changes by a few favourite demo questions.
- Using an LLM judge without calibration.
- Letting the golden set go stale while real usage evolves.
- Only engineers looking at eval results; PMs and domain experts must read failures too.
How to measure success
Every release has an eval report; scores trend up by slice; the golden set grows from real failures; and nobody ships a change the evals have not seen.
Video lecture: Evals-driven development with your engineering team
Lecture coming soon · 16 chapters · about 9 minutes. Read the full transcript below.
- Evals-driven development
- Analogy: marathon training
- The loop
- Error analysis
- Simple example: 20 failures
- Five eval types
- Evals in the workflow
- Business example: KSA HR (illustrative)
- Hands-on: eval spec (YAML)
- Common mistakes
- Another example: UK insurance broker
- The weekly eval hour
- FAQ: do evals slow us down?
- Try this now
- Watch me do it
- Recap
Lecture transcript
Evals-driven development
In most software teams, the question is the feature done has a clear answer: the tests pass. In AI teams, without evaluations, the answer is usually the demo looked good, which is not an answer at all. In this lesson you will learn evals-driven development: the loop that makes AI progress measurable, the error analysis step most teams skip, the types of evals and who owns them, and a shared eval spec for PMs and engineers.
Analogy: marathon training
Here is an analogy. Developing AI without evals is like training for a marathon without ever timing your runs. You feel faster some days and slower on others, but you cannot tell whether a new shoe or a new diet helped. Evals are the stopwatch and the route you run every week. Same route, same stopwatch, so you can see whether a change actually made you faster, and on which part of the course.
The loop
The loop has seven steps. Define the golden set, rubric and thresholds from your PRD. Baseline the current system. Analyse errors by reading failures and grouping them. Pick the biggest failure category and a hypothesis. Make one change: prompt, retrieval, model, tool, data or UX. Re-run the evals and compare with the baseline, checking slices and regressions. Then ship or iterate, and add new real-world failures to the golden set.
Error analysis
Now the step teams skip: error analysis. Scores tell you how much is wrong. Reading failures tells you what is wrong. Take fifty failures and tag each with a category: retrieval missed the document, right documents but wrong reasoning, outdated source, tone or format, should have abstained, language issues. Count them. Suddenly your roadmap is obvious: fix the biggest category first. PMs are often the best people to lead this, because they understand the users and the domain.
Simple example: 20 failures
A simple example. A small e-commerce assistant fails on twenty questions. The PM reads them all. Twelve are about delivery times to specific cities, and the assistant keeps citing the general policy. Five are tone issues. Three are about products that no longer exist. The fix for the biggest category is not a new model. It is adding the city-level delivery table to the knowledge base. After one change, twelve failures become two.
Five eval types
There are five kinds of evals, and ownership matters. Code-based checks, like valid JSON, required fields, length and citations present: fast and cheap, owned by engineers. Reference checks against expected answers: shared. Human rubric scoring by domain experts: the PM organises it. LLM-as-judge, where a model scores outputs with your rubric: useful for scale, but only after calibration against human scores. And online signals, like feedback, edits and escalations: owned by the PM and analytics.
Evals in the workflow
Put evals into the workflow. Fast checks and a small subset on every change. The full golden set before every release and whenever the model or provider changes. A dashboard showing scores over time by slice, plus cost and latency. And regression gates: a release is blocked if any slice drops beyond an agreed tolerance. Tools like promptfoo, Inspect and observability platforms help, but the discipline matters more than the tool.
Business example: KSA HR (illustrative)
Now a realistic example over four weeks, with illustrative numbers. A Saudi HR assistant starts with a rubric mean of three point six. Error analysis shows forty per cent of failures are retrieval misses on Arabic terms and abbreviations, like EOS for end-of-service benefits. Week two adds Arabic synonyms and hybrid search: three point nine. Week three removes outdated policies: four point one. Week four tightens abstention, and unanswerable accuracy jumps from seventy-two to ninety-one per cent. The launch gate is met. After launch, thirty real failures join the golden set.
Hands-on: eval spec (YAML)
The hands-on is a shared eval spec in YAML. It names the golden set file and its tags, lists checks with their type and threshold: JSON validity, citations, a calibrated rubric judge, and abstention on unanswerable questions. It names the slices, the regression gate, when each suite runs, and who owns what. PMs and engineers can both read it, and it lives in the repository next to the code.
Common mistakes
Common mistakes. Judging changes by a few favourite demo questions. Trusting an LLM judge that nobody calibrated. Letting the golden set go stale while real usage moves on. And leaving eval results to engineers alone. PMs and domain experts must read failures too, because they know what a real user would consider wrong.
Another example: UK insurance broker
Another example. A UK insurance broker's claims assistant kept failing on questions about policy excess. Error analysis showed the problem was not the model at all: the knowledge base had three versions of the excess table, and retrieval picked the oldest one a third of the time. Removing old versions and adding an effective-date filter fixed most of the failures. No prompt change, no model change. Reading the failures pointed straight to the fix.
The weekly eval hour
How big should the team's weekly eval ritual be? Small. One hour, three people: the PM, an engineer and a domain expert. Look at the latest scores by slice, read ten to twenty fresh failures together, agree on the biggest category, and decide one change for the coming week. Consistency beats intensity. Teams that do this every week improve steadily; teams that do a big evaluation push once a quarter drift in between.
FAQ: do evals slow us down?
A question from engineers: won't evaluation slow us down? At first, a little. Within weeks, it speeds you up, because you stop debating opinions and stop re-breaking things you already fixed. A fast smoke subset on every change takes minutes. The full suite before release takes an hour or two. Compare that with a week of firefighting after a silent regression reaches customers. Evals are the fastest way to move safely.
Try this now
Try this now. Take thirty failures from any AI feature you work with, from feedback, logs or your own testing. Tag each one with a category, count them, and put the counts in a table. Then write one sentence: the biggest category, the fix you will try, and the eval result that would prove it worked.
Watch me do it
Watch me do it. Our HR assistant scores three point eight on the golden set, below the four point two gate. Instead of tweaking the prompt, I run error analysis. I export the forty lowest-scoring answers and, with an HR specialist, tag each one. Seventeen: retrieval missed the document, mostly questions using the abbreviation EOS for end-of-service benefits. Nine: an outdated policy was cited. Eight: too long for mobile. Six: should have said not found. I bring the counts to the engineer. We add abbreviations and Arabic synonyms to the index metadata and rerun: three point nine nine. Next week, we remove outdated policies: four point one. Then a length limit and a stronger abstention rule: four point two five, with no slice regressing more than the tolerance. The eval spec runs in the pipeline, so every one of those changes had a report attached.
Recap
Recap. Evals are your specification, test suite and progress report. Run the loop, and never skip error analysis. Use the right mix of code checks, references, human rubrics, calibrated judges and online signals. Gate releases on slice regressions, and grow the golden set from real failures. Next: launching safely with staged rollouts and safety reviews.
Key takeaways
- Evals are the spec, test suite and progress report for AI features
- Run the loop: define, baseline, analyse errors, fix the biggest category, re-run, ship
- Error analysis turns scores into a roadmap; PMs should lead it
- Combine code checks, reference checks, human rubrics, calibrated LLM judges and online signals
- Gate releases on slice regressions and grow the golden set from real failures
Try it
Categorise 30 failures from an AI feature into an error-analysis table, pick the biggest category, propose a fix and define how the eval will show it worked.