Building AI Products & Workflows · Finding and scoping high-value use cases · lesson 2 of 18 · 11 min
Scoping and defining success
A use case is not a scope
"AI for customer support" is a theme. A scope is: "Draft first replies to order-status and returns emails in English and Arabic, for human agents to review and send, reducing median handling time without lowering quality scores." Clear scope is what turns ideas into shippable products.
The one-page AI brief
Before building or buying, write a one-page brief:
Problem: What pain, for whom, how often?
Users: Who uses the output? Who is affected by it?
Scope (in): Which inputs, tasks, languages, channels?
Scope (out): What we explicitly won't do in v1
AI's role: Draft / recommend / classify / act (with what approval?)
Success: Primary metric + target; quality guardrail metrics
Risks: What happens if it's wrong? Who could be harmed?
Data: What data it needs, where it lives, sensitivity
Evaluation: How we'll test before launch
Owner: Accountable person
Choosing the AI's role
The same use case can be designed with very different levels of autonomy:
- Assist: AI suggests; human decides and acts (drafts, summaries, recommendations).
- Automate with review: AI acts, a human approves before it takes effect.
- Automate with audit: AI acts; humans review samples afterwards.
- Full automation: AI acts; exceptions escalate.
Choose based on the cost of errors, the ability to detect errors, volume, and user trust. Most first versions should be at level 1 or 2.
Success metrics: three layers
- Business outcome: the reason the project exists (handling time, conversion, cost per case, revenue, customer satisfaction).
- Quality metrics: accuracy, groundedness, policy compliance, tone, measured by evaluation and review.
- Adoption and experience: usage rate, edit rate (how much humans change AI drafts), user satisfaction, time to complete.
Plus guardrail metrics that must not get worse: complaint rate, error escalations, fairness indicators across user groups.
Set targets before launch. "We'll see how it goes" makes every result look like success.
Establish a baseline
You can only claim improvement against a baseline. Before launch, measure the current process: time per task, error rate, customer satisfaction, cost. If the baseline doesn't exist, spend a week or two collecting it. It will be the most persuasive number in your launch review.
Define "good enough"
AI outputs are probabilistic; perfection is not the bar. Define acceptable quality by comparing with the current human process and the cost of errors:
- For internal drafts reviewed by experts, moderate accuracy can still save a lot of time.
- For customer-facing automated replies, the bar is much higher, and some categories may be excluded entirely.
Worked example
An insurance broker scopes a document assistant:
- Problem: brokers spend significant time extracting key terms from policy documents to compare quotes.
- Scope in: extracting 15 defined fields from PDF policy schedules from the top 10 insurers; English only in v1.
- Scope out: giving advice to customers; handwritten documents; other languages.
- AI's role: extract fields with page citations; broker reviews in a side-by-side view.
- Success: median extraction time per document reduced substantially; field accuracy above an agreed threshold on the evaluation set; broker satisfaction above baseline survey.
- Guardrails: no increase in quote errors reported by customers.
- Owner: head of broking operations.
This brief makes the build smaller, the evaluation concrete, and the launch decision clear.
Hands-on: the AI brief as a living config file
Keep the brief next to the code or workflow, in a format both humans and tools can read. Teams increasingly store it as YAML so evaluation harnesses and dashboards can pick up the thresholds automatically.
# ai-brief.yaml (owner: head of broking operations)
feature: policy-schedule-extractor
version: 1
problem: Brokers spend significant time extracting key terms from policy schedules to compare quotes.
users: [brokers]
scope_in:
documents: PDF policy schedules from our top 10 insurers
fields: 15 defined fields (see fields.yaml)
languages: [en]
scope_out: [customer advice, handwritten documents, non-English documents]
ai_role: extract_with_review # assist | extract_with_review | automate_with_audit | full_automation
success:
primary: {metric: median_minutes_per_document, baseline: measure_in_week_0, target: "reduce substantially vs baseline"}
quality:
- {metric: field_accuracy, threshold: 0.95, measured_on: eval_set_v1}
- {metric: citation_present_rate, threshold: 0.99}
adoption: [{metric: weekly_active_brokers_share}, {metric: fields_edited_per_document}]
guardrails:
- {metric: customer_reported_quote_errors, rule: "no increase vs baseline"}
kill_criteria:
- field_accuracy below threshold after two improvement iterations
- fields_edited_per_document shows no time saving after 6 weeks of pilot
risks: [wrong field values in quotes, confidential data exposure]
data: {sensitivity: confidential, location: document store EU region, retention_days: 30}
evaluation: {set: eval_set_v1, size: 120, graders: [exact_match, human_review_sample]}
Then stress-test the brief with an AI reviewer before anyone builds:
Act as a sceptical product reviewer. Read the AI brief below and list:
1. Ambiguities that would let two engineers build different things.
2. Metrics without a baseline, owner or measurement method.
3. Risks that the scope or the AI role does not address.
4. What a kill-criteria review would most likely find at week 6.
Be specific; quote the lines you are criticising.
<brief>{{PASTE ai-brief.yaml}}</brief>
Fix what it finds, get the owner's sign-off, and record a baseline before the first prototype runs.
Going further
Write "kill criteria" as well as success criteria: the conditions under which you'll stop or pivot (for example, accuracy plateaus below the threshold after two iterations, or edit rates show drafts aren't saving time). Pre-committing prevents sunk-cost projects that linger for years.
Video lecture: Scoping and defining success
Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.
- Scoping AI work
- Analogy: 'make the house better'
- The one-page AI brief
- Choosing the AI's role
- Three layers + guardrails
- Baselines and 'good enough'
- Simple example: café reviews
- Worked example: insurance broker
- Business example (illustrative)
- Kill criteria + hands-on
- Common mistakes
- How you'll know it worked
- Watch me do it: ai-brief.yaml
- Recap
- Try this now (30 minutes)
Lecture transcript
Scoping AI work
AI for customer support is not a project. It's a theme. And themes are where budgets go to die, because nobody can say when they're done or whether they worked. In this lesson you'll turn a theme into a scope, choose how much autonomy the AI should have, define success in three layers plus guardrails, and write kill criteria before you start. You'll leave with a one-page brief that makes the build smaller and the launch decision obvious.
Analogy: 'make the house better'
Here's an analogy. Starting an AI project without a scope is like hiring a builder and saying, make the house better. You'll get something, you'll pay for it, and you won't agree on whether it's finished. A scope is the drawing with measurements. It tells everyone what's being built, what isn't, and how you'll know it's done. It feels slower on day one and saves weeks by day thirty.
The one-page AI brief
Compare the theme with this scope. Draft first replies to order-status and returns emails, in English and Arabic, for human agents to review and send, reducing median handling time without lowering quality scores. Now you know the inputs, the languages, the channel, who's in control, and what success means. Write every AI project as a one-page brief: problem, users, what's in scope and explicitly out, the AI's role, success metrics, risks, data, how you'll evaluate, and a single accountable owner.
Choosing the AI's role
Next, choose the AI's role. The same use case can run at four levels of autonomy. Assist: AI suggests, and a human decides and acts. Automate with review: AI acts, but a human approves before it takes effect. Automate with audit: AI acts, and humans review samples afterwards. Full automation: AI acts and only exceptions escalate. Choose based on the cost of errors, how easily errors are spotted, volume and user trust. Most first versions belong at assist or automate with review.
Three layers + guardrails
Define success in three layers. The business outcome is why the project exists: handling time, conversion, cost per case, revenue or satisfaction. Quality metrics cover accuracy, groundedness, policy compliance and tone, measured by evaluation and review. Adoption and experience cover usage, how much people edit AI drafts, satisfaction and time to complete. Then add guardrail metrics that must not get worse, like complaint rates, escalations and fairness across user groups. Set targets before launch. We'll see how it goes makes every result look like success.
Baselines and 'good enough'
You can only claim improvement against a baseline. Before launch, measure today's process: time per task, error rate, satisfaction, cost. If the numbers don't exist, spend a week or two collecting them; they'll be the most persuasive figures in your launch review. And define good enough by comparing with the current human process and the cost of errors. Internal drafts reviewed by experts can be useful at moderate accuracy. Customer-facing automated replies need a much higher bar, and some categories should be excluded entirely.
Simple example: café reviews
A simple example before the insurance one. A café chain wants AI for customer reviews. Theme: use AI on reviews. Scope: every Monday, summarise last week's Google and delivery-app reviews per branch into three positives, three complaints and one suggested action, for branch managers to read. Out of scope: replying to reviews. Success: managers read it and act on at least one item a week. One paragraph, and suddenly everyone knows what to build.
Worked example: insurance broker
Here's a brief in action. An insurance broker scopes a document assistant. Problem: brokers spend significant time extracting key terms from policy schedules. In scope: fifteen defined fields from PDF schedules from their top ten insurers, English only. Out of scope: advice to customers, handwritten documents, other languages. AI's role: extract with page citations while the broker reviews side by side. Success: faster extraction, field accuracy above an agreed threshold, and better broker satisfaction. Guardrail: no increase in quote errors. Owner: head of broking operations.
Business example (illustrative)
Illustrative numbers for the broker. Baseline measurement showed about thirty-five minutes per policy schedule and a few quote errors a month traced to extraction. The target was under fifteen minutes with field accuracy of at least ninety-five percent on a hundred and twenty test documents. The pilot reached about twelve minutes and ninety-six percent. Because the targets were written down first, the launch decision took one meeting.
Kill criteria + hands-on
Now write kill criteria as well as success criteria: the conditions under which you'll stop or pivot. For example, accuracy plateaus below the threshold after two improvement rounds, or edit rates show drafts aren't saving time after six weeks. Pre-committing prevents zombie projects. In the hands-on section you'll store the brief as a YAML file your evaluation harness and dashboards can read, and use a sceptical AI reviewer prompt that hunts for ambiguities, metrics without baselines, and unaddressed risks before anyone builds.
Common mistakes
Common scoping mistakes. Leaving out the out-of-scope list, so the project grows every week. Choosing only adoption metrics, which can rise while quality falls. Setting targets after launch. Forgetting guardrail metrics, so a faster process quietly produces more complaints. And naming a committee as owner. If nobody's name is on it, nobody will stop it when the kill criteria are met.
How you'll know it worked
How will you know your scoping worked? Engineers and designers can estimate the work from the brief without a meeting. The evaluation set follows directly from the scope. At the launch review, the decision is quick, because targets and kill criteria were agreed in advance. And six weeks in, nobody is arguing about what the project was supposed to do.
Watch me do it: ai-brief.yaml
Watch me do it. I open ai-brief dot yaml. First, feature, version and owner. Next, problem and users in one line each. Then scope in: PDF schedules from our top ten insurers, fifteen fields, English. Scope out: advice, handwritten documents, other languages. AI role: extract with review. Success: the primary metric is median minutes per document, with the baseline measured in week zero; quality thresholds for field accuracy and citation rate; and two adoption metrics. Guardrails: no increase in customer-reported quote errors. Kill criteria: accuracy below threshold after two iterations, or no time saving after six weeks. Then I paste the file into the sceptical reviewer prompt. It flags three things: no definition of how minutes are measured, no owner for the evaluation set, and no plan for insurer format changes. I add a timing method, name an evaluation owner and add a monthly format check.
Recap
To recap: turn themes into scopes with clear in and out boundaries. Choose the AI's autonomy by error cost, detectability, volume and trust. Define business, quality, adoption and guardrail metrics, measure a baseline and write kill criteria. Your next step is to write a one-page brief for your top use case, including out-of-scope items and kill criteria, and run it past the sceptical reviewer prompt. Next module, we'll prototype it cheaply on real data.
Try this now (30 minutes)
Try this now. Take your top use case from the last lesson and fill in the one-page brief: problem, users, scope in, scope out, the AI's role, three layers of metrics, a guardrail, kill criteria and one named owner. Then paste it into the sceptical reviewer prompt and fix the three most important issues it finds. Thirty minutes, and you'll have a brief you can actually build from.
Key takeaways
- Turn themes into scopes with explicit in-scope and out-of-scope boundaries.
- Write a one-page AI brief: problem, users, scope, AI's role, success, risks, data, evaluation and owner.
- Choose the AI's autonomy level based on error cost, detectability, volume and trust.
- Define business, quality, adoption and guardrail metrics; measure a baseline and set kill criteria.
Try it
Write a one-page AI brief for your top-priority use case, including out-of-scope items, the AI's autonomy level, three layers of metrics, a baseline plan and kill criteria.