AI Product Management: From Idea to Reliable AI Features · Specs, prototypes and sourcing decisions · lesson 6 of 16 · 16 min
Writing AI PRDs with evaluation criteria
Why AI PRDs are different
A traditional PRD says what the feature does. For deterministic software, if engineering builds to spec, it works. For AI features, "builds to spec" is not enough: the same code produces different outputs on different inputs, and quality is a distribution, not a checkbox. An AI PRD must therefore define what good looks like, how it will be measured, and what happens when it is not good.
The AI PRD skeleton
- Problem and user: who, what job, current pain, evidence (tickets, interviews, data).
- Outcome metrics: the business and user outcomes you expect to move (for example first-response time, resolution rate, conversion).
- Scope and autonomy level: which actions the AI takes, at which autonomy level (inform, suggest, draft, approve, act and notify), with limits.
- Inputs and context: what data the AI can access (documents, CRM fields, history), permissions, languages, channels.
- Output specification: format, tone, length, structure, citations, what must never appear.
- Quality bar and evaluation plan: the eval set, metrics, rubric, thresholds for launch, slices that must not regress.
- Failure modes and mitigations: from your failure-mode register.
- UX requirements: streaming, citations, not-found states, undo, handoff, disclosure, feedback.
- Cost and latency budgets: target cost per task and p95 latency.
- Launch plan: stages, success gates, kill switch, monitoring.
- Open questions and risks: including legal, privacy and regulatory checks.
Writing evaluation criteria engineers can build against
Vague: "Answers should be accurate and helpful."
Buildable:
- Correctness: ≥ 90% of answers on the 200-question golden set are fully correct per the rubric; 0 answers contradict cited sources.
- Grounding: 100% of factual claims have a citation to a passage that supports them (spot-checked on 50 answers).
- Abstention: ≥ 90% of the 30 unanswerable questions receive the "not found" response.
- Slices: Arabic and Urdu slices each within 5 points of English.
- Safety: 0 failures on the 40-case red-team set.
- Latency: p95 time to first token ≤ 1.5 s; p95 full answer ≤ 8 s.
- Cost: ≤ $0.02 per answered question at expected volume (illustrative budget).
Each criterion has a dataset, a measurement method and a threshold. That is what makes it testable.
The golden set: the PM's most important artefact
The golden set (or eval set) is a collection of realistic inputs with expected outputs or grading rubrics. As PM you own its coverage: common cases, edge cases, languages, user types and known failure modes. Engineers and domain experts help build it. A practical start: 50–100 examples drawn from real data, grown to a few hundred as you learn. Version it like code.
A rubric example for a support reply:
| Score | Meaning | |---|---| | 5 | Correct, complete, cites policy, right tone, clear next step | | 4 | Correct and complete, minor tone or formatting issue | | 3 | Mostly correct, missing a useful detail | | 2 | Partly wrong or missing key information | | 1 | Wrong, unsafe or off-policy |
Hands-on: AI PRD template (quality section)
## Quality bar and evaluation plan
**Golden set:** v1.2, 220 cases (160 answerable EN/AR/UR, 30 unanswerable, 30 red-team). Owner: PM. Location: /evals/support-assistant/
**Metrics and launch thresholds:**
| Metric | Method | Threshold | Current |
|---|---|---|---|
| Rubric score (1-5) | 2 trained reviewers, disagreements adjudicated | mean ≥ 4.2, ≤ 3% scored 1 | 3.9 |
| Grounded claims | Spot-check 50 answers | 100% | 96% |
| Abstention | Unanswerable subset | ≥ 90% | 83% |
| Slice gap (AR/UR vs EN) | Rubric mean difference | ≤ 0.3 | 0.5 (UR) |
| Red-team failures | 30-case set | 0 critical | 1 critical |
| p95 first token | Load test at 5 req/s | ≤ 1.5 s | 1.2 s |
| Cost per answer | Token logs x current prices | ≤ budget | within |
**Launch decision:** all thresholds met, or explicit sign-off with mitigation for any exception.
Worked example: a WhatsApp sales assistant for a UAE car dealer
The first PRD draft said "assistant answers customer questions and books test drives". The revised PRD specified: autonomy (answers from the approved inventory and pricing feed; books test drives with confirmation; never negotiates price), languages (English and Arabic), outputs (short messages, one question at a time, price only from the feed with a timestamp), evaluation (150-conversation golden set covering inventory questions, financing questions routed to humans, Arabic dialect variation, off-topic and abusive messages), thresholds, a handoff rule (any financing or trade-in question goes to a salesperson within business hours), and a cost budget per conversation. Engineering estimated faster, legal approved faster, and launch criteria were unambiguous.
Pitfalls
- PRDs without evaluation criteria ("we'll know it when we see it").
- Thresholds set after seeing results.
- Golden sets built only from easy, English examples.
- No cost or latency budget, so the "best" solution is unaffordable.
How to measure success
Engineers can build and test against your PRD without asking what "good" means; every criterion has data, a method and a threshold; and launch decisions follow the thresholds.
Video lecture: Writing AI PRDs with evaluation criteria
Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.
- Writing AI PRDs
- Why AI PRDs differ
- AI PRD skeleton
- Vague vs buildable
- The golden set
- A 5-point rubric
- Simple example: language school
- Business example: UAE car dealer
- Hands-on: quality section template
- Common mistakes
- Another example: local council assistant
- Two quick PRD reviews
- Try this now
- Watch me do it
- Recap
Lecture transcript
Writing AI PRDs
A product manager hands engineering a PRD that says: the assistant should answer customer questions accurately and helpfully. Six weeks later, the demo looks great, and nobody can say whether it is ready to launch. Is eighty per cent accurate good enough? On which questions? In which languages? In this lesson you will learn to write AI PRDs that answer those questions up front, with evaluation criteria engineers can actually build and test against.
Why AI PRDs differ
Why are AI PRDs different? For normal software, if engineering builds to spec, it works. For AI, the same code gives different outputs on different inputs, and quality is a distribution, not a checkbox. Here is an analogy. A traditional PRD is like a blueprint for a chair: build it to these measurements and it is done. An AI PRD is more like hiring criteria for a customer service agent: you describe the job, then define how you will assess performance across many real conversations.
AI PRD skeleton
Here is the skeleton, eleven sections. Problem and user. Outcome metrics. Scope and autonomy level for each action. Inputs and context, including permissions, languages and channels. The output specification: format, tone, length, citations and what must never appear. The quality bar and evaluation plan. Failure modes and mitigations. UX requirements. Cost and latency budgets. The launch plan with gates and a kill switch. And open questions, including legal and privacy checks.
Vague vs buildable
Now the heart of it: evaluation criteria. Compare two versions. Vague: answers should be accurate and helpful. Buildable: at least ninety per cent of answers on the two-hundred-question golden set are fully correct according to the rubric, and zero contradict their cited sources. At least ninety per cent of the thirty unanswerable questions get the not-found response. Arabic and Urdu slices within five points of English. Zero failures on the red-team set. Every buildable criterion has three parts: a dataset, a measurement method and a threshold.
The golden set
Which brings us to the golden set, the most important artefact a PM owns in AI work. It is a collection of realistic inputs with expected answers or grading rubrics. You own its coverage: common cases, edge cases, languages, user types and known failure modes. Start with fifty to a hundred real examples and grow to a few hundred as you learn. Version it like code. When someone asks whether the new model is better, the golden set answers in an afternoon.
A 5-point rubric
For open-ended outputs, you need a rubric. A simple five-point rubric for support replies: five means correct, complete, cites policy, right tone and a clear next step. Four is correct and complete with a minor tone or formatting issue. Three is mostly correct but missing a useful detail. Two is partly wrong or missing key information. One is wrong, unsafe or off-policy. Train two reviewers on it with examples, and measure how often they agree.
Simple example: language school
A simple example. A small language school wants an AI to answer questions about course dates and fees. Their mini PRD quality section has four lines: a golden set of forty real questions from last term; at least thirty-eight must be answered correctly from the course catalogue with a citation; all five questions about topics not in the catalogue must get a polite not-found with the school's phone number; and answers must arrive within three seconds. Four lines, and everyone knows when it is ready.
Business example: UAE car dealer
Now a realistic business example. A UAE car dealer wanted a WhatsApp sales assistant. The first draft said: answers customer questions and books test drives. The revised PRD said: answers only from the approved inventory and pricing feed, with a timestamp on prices; books test drives with customer confirmation; never negotiates price; routes any financing or trade-in question to a salesperson within business hours; works in English and Arabic including dialect variation. It had a hundred-and-fifty-conversation golden set, thresholds and a cost budget. Engineering estimated faster and legal approved faster.
Hands-on: quality section template
The hands-on template in the lesson is a quality section you can paste into any PRD. It names the golden set version and composition, then a table of metrics with the method, threshold and current value for each: rubric score, grounded claims, abstention, slice gaps, red-team failures, first-token latency and cost per answer. And it ends with a launch decision rule: all thresholds met, or explicit sign-off with a mitigation for any exception.
Common mistakes
Common mistakes. PRDs with no evaluation criteria, where the plan is we will know it when we see it. Thresholds invented after seeing results, which turns evaluation into theatre. Golden sets built only from easy English examples. And no cost or latency budget, so the best solution turns out to be unaffordable at real volume.
Another example: local council assistant
Another example, from the UK public sector style of work. A local council wants an AI assistant to help residents find the right service page, like bin collections or council tax. The PRD specifies: answers only from the council's published pages with links; never gives legal or benefits eligibility decisions; offers a phone number and opening hours when unsure; plain English at a reading level most residents can follow; and accessibility requirements. The golden set includes questions from real resident emails, including angry and confused ones.
Two quick PRD reviews
One practical tip for PRD reviews. Ask engineering to read your quality section first and answer one question: could you write an automated test for each line? Any line they cannot test needs rewriting. Then ask a domain expert to read your golden set and answer: are these the questions real users ask, and are these the answers you would give? Those two reviews catch most PRD problems in under an hour.
Try this now
Try this now. Take one AI feature, real or planned, and write three buildable criteria for it. Each must name a dataset, a measurement method and a threshold. Then write ten real inputs for its golden set, including at least two edge cases and one input in a second language your users speak. You will immediately see where the feature definition is still fuzzy.
Watch me do it
Watch me do it. I'm rewriting the quality section of a PRD for a meeting-summary feature. The original line says summaries should be accurate and useful. I replace it with four criteria. First, key decisions: ninety-five per cent of decisions recorded in the meeting appear in the summary, measured on a golden set of forty real meetings with human-written decision lists. Second, no invented action items: zero on the golden set. Third, Arabic meetings score within point three of English on the rubric. Fourth, cost under a set budget per hour of audio. For each I add the method and who measures it. Then I build the golden set composition: thirty English, ten Arabic, including five noisy recordings. Finally I ask an engineer the key question: can you write a test for each line? She says yes for three, and asks how we will define a decision. I add a definition. Now it is buildable.
Recap
Recap. AI PRDs define what good looks like, how you will measure it and what happens when it is not good. Use the eleven-section skeleton. Write criteria with a dataset, method and threshold. Own the golden set and its coverage. And set thresholds before you see results. In the next lesson, we will prototype quickly to test these ideas before engineering builds anything.
Key takeaways
- AI PRDs must define what good looks like, how it is measured, and what happens when it is not good
- Include autonomy level, inputs, output spec, quality bar, failure modes, UX, cost/latency budgets and launch plan
- Write buildable criteria: dataset + method + threshold
- The PM owns the golden set's coverage; version it like code
- Set thresholds before seeing results
Try it
Rewrite the quality section of an AI feature PRD using the template: golden set, metrics, methods, thresholds and a launch decision rule.