---
title: "Where AI creates value (and where it does not)"
description: "The PM's first job: pick the right problems Most failed AI features were never going to work. Not because the model was weak, but because the problem did…"
url: https://optimizeall.com/learn/ai-product-management/where-ai-creates-value
updated: 2026-10-05
---

AI Product Management: From Idea to Reliable AI Features · AI product strategy · lesson 1 of 16 · 14 min

# Where AI creates value (and where it does not)

## The PM's first job: pick the right problems

Most failed AI features were never going to work. Not because the model was weak, but because the problem did not suit AI, the value was small, or the cost of mistakes was too high. Great AI product managers spend more time choosing problems than choosing models.

## The value map: four kinds of AI value

1. **Doing existing work faster or cheaper** (efficiency): drafting replies, summarising calls, classifying tickets, extracting invoice fields.
2. **Doing work that was not economical before** (new scale): personalised onboarding for every customer, reviewing every contract instead of a sample, translating every help article into Arabic and Urdu.
3. **Better decisions** (quality): surfacing risks in documents, recommending next best actions, spotting anomalies.
4. **New experiences** (differentiation): conversational interfaces, AI agents that complete tasks, generative creative tools.

Efficiency is easiest to measure; new scale and new experiences are often where the biggest wins are.

## The AI fit test

Score a candidate problem 1–5 on each dimension:

| Dimension | Good fit looks like |
|---|---|
| **Language or perception heavy** | Work involves reading, writing, summarising, classifying, seeing or hearing |
| **Tolerance for imperfection** | Occasional errors are cheap, catchable or reviewable |
| **Verifiability** | A human or program can check outputs quickly |
| **Frequency and volume** | The task happens often enough to matter |
| **Data and context access** | The AI can get the information it needs (documents, systems, history) |
| **Clear success measure** | You can say what "good" means and measure it |

High scores on tolerance and verifiability matter most. A task where mistakes are expensive **and** hard to spot (for example, final medical dosing decisions or legal advice sent directly to consumers) is a poor candidate for full automation, though it may still suit assistance with expert review.

## Where AI usually does not create value

- **Deterministic logic** that plain code handles perfectly (tax rules, price calculations). Use code; let AI handle the messy language around it.
- **Low-frequency tasks** where building and evaluating costs more than it saves.
- **Problems that are really process or data problems.** If nobody agrees what the right answer is, AI will not settle it.
- **"AI for AI's sake"**: features added for a press release that do not change a user outcome.

## Sizing the opportunity

A simple formula for efficiency value:

```text
annual value ≈ task volume per year × minutes saved per task × share of tasks the AI handles well × loaded cost per minute
```

Example (illustrative): a Karachi insurance broker handles 60,000 claim emails a year. Drafting a first reply takes 6 minutes; with AI drafts reviewed by agents it takes 2.5 minutes, on the 80% of emails where drafts are usable. Saving: 60,000 × 3.5 × 0.8 = 168,000 minutes (2,800 hours) a year. Multiply by loaded cost per hour to compare against build and run costs. Then ask what else changes: faster replies may improve retention, which may be worth more than the time saved.

## Hands-on: an opportunity scorecard

```markdown
| Opportunity | Value type | Volume/yr | Fit score (6-30) | Est. annual value | Risk of errors | Data access | Priority |
|---|---|---|---|---|---|---|---|
| Draft replies to claim emails | Efficiency | 60,000 | 25 | 2,800 agent hours | Low (reviewed) | Email + policy DB | High |
| Auto-approve small claims | Efficiency | 8,000 | 14 | High | High (money, fairness) | Claims system | Later, with controls |
| Arabic/Urdu policy explainer | New scale | 20,000 views | 22 | Conversion, fewer calls | Medium | Policy docs | Medium |
```

Run a 60-minute workshop with frontline staff: list 20 candidate tasks, score them together, and pick two for discovery. Frontline staff know where time goes and where mistakes hurt.

## Worked example: a creator agency in Dubai

An agency managing 30 creators lists candidates: caption drafting, brand-brief summarisation, sponsorship contract review, audience-comment triage and performance report writing. Scoring shows comment triage (huge volume, low error cost, easy to verify) and report writing (weekly, time-consuming, reviewable) as top candidates. Contract review scores high value but also high risk; they scope it as "highlight clauses for the lawyer" rather than "approve contracts".

## Pitfalls

- Starting from a model capability ("we have an agent framework") rather than a user problem.
- Counting minutes saved without checking whether the saved time creates value.
- Ignoring the cost of reviewing AI output, which can erase savings.
- Choosing a high-risk task as the first AI feature.

## How to measure success

A scored, prioritised opportunity list with value estimates, error-cost notes and data availability, agreed with the people who do the work.

## Video lecture: Where AI creates value (and where it does not)

Lecture coming soon · 16 chapters · about 9 minutes. Read the full transcript below.

1. Where AI creates value
2. Analogy: a tireless, well-read intern
3. Four kinds of value
4. The fit test (1–5 each)
5. Where AI does not help
6. Sizing (illustrative)
7. Net out review time
8. Hands-on: scorecard workshop
9. Simple example: Lahore tailoring shop
10. Second example: Jeddah logistics (illustrative)
11. Worked example: Dubai creator agency
12. FAQ: experiment everywhere or big bets?
13. FAQ: measuring quality value
14. Try this now
15. Watch me do it
16. Recap

## Lecture transcript

### Where AI creates value

Here is a pattern I have seen at startups and enterprises alike. A team ships an impressive AI feature. Usage spikes for two weeks, then flatlines. The model was fine. The problem was wrong. In this lesson you will learn where AI actually creates value, a simple fit test for candidate problems, how to size the opportunity, and a scorecard you can run with your team this week.

### Analogy: a tireless, well-read intern

Here is an analogy that helps. Think of AI like hiring a very fast, very well-read intern who never gets tired but occasionally gets things confidently wrong. Where would you put that intern to work? Not signing contracts on their own. Not doing sums a calculator does perfectly. You would give them high-volume drafting, sorting and summarising, where a quick check catches their mistakes. That instinct is the whole fit test in one picture: lots of language work, cheap mistakes, easy checking.

### Four kinds of value

AI creates value in four ways. Efficiency: doing existing work faster or cheaper, like drafting replies or extracting invoice fields. New scale: work that was not economical before, like reviewing every contract instead of a sample, or translating every help article into Arabic and Urdu. Better decisions: surfacing risks and anomalies. And new experiences: conversational interfaces, agents that complete tasks, creative tools. Efficiency is easiest to measure, but new scale and new experiences often hold the biggest wins.

### The fit test (1–5 each)

Now the fit test. Score each candidate from one to five on six dimensions. Is the work heavy on language or perception? Can you tolerate occasional mistakes? Can a person or program check outputs quickly? Does it happen often? Can the AI access the data it needs? And can you say what good means? Tolerance and verifiability matter most. A task where mistakes are both expensive and hard to spot is a poor candidate for automation, though it may still suit assistance with expert review.

### Where AI does not help

Just as important is knowing where AI does not help. Deterministic logic, like tax rules and price calculations, belongs in plain code. Rare tasks cost more to build and evaluate than they save. Process problems, where people disagree on the right answer, will not be solved by a model. And features added for a press release, that do not change a user outcome, are expensive decorations.

### Sizing (illustrative)

To size an efficiency opportunity: task volume per year, times minutes saved per task, times the share of tasks AI handles well, times the loaded cost per minute. An illustrative example: a Karachi insurance broker gets sixty thousand claim emails a year. A first reply takes six minutes by hand, two and a half with reviewed AI drafts, and drafts are usable eighty per cent of the time. That is about two thousand eight hundred agent hours a year. Then ask what else changes, because faster replies might improve retention.

### Net out review time

Remember to net out review time. If agents spend three minutes checking a draft that saves four, your real saving is one minute. Review cost is the silent killer of AI business cases. Design so that checking is fast: show sources, highlight uncertain parts, and make edits easy. We will come back to these patterns in the user experience module.

### Hands-on: scorecard workshop

Your hands-on tool is an opportunity scorecard: opportunity, value type, yearly volume, fit score, estimated value, error risk, data access and priority. Run a sixty-minute workshop with frontline staff. List ten to twenty candidate tasks, score them together, and pick two for discovery. Frontline people know where time goes and where mistakes hurt, and involving them builds adoption from day one.

### Simple example: Lahore tailoring shop

A simple example. A tailoring shop in Lahore gets dozens of WhatsApp messages a day asking about prices, delivery times and fabric options. The owner scores the task: heavy on language, yes. Mistakes cheap, yes, because the owner reads replies before sending. Easy to verify, yes. Frequent, yes, every day. Data access: a one-page price list and fabric catalogue. Clear success: fewer minutes per message. It scores high on every dimension, so AI-drafted replies are a perfect first feature.

### Second example: Jeddah logistics (illustrative)

A second, bigger example with illustrative numbers. A logistics company in Jeddah looks at two ideas. Idea one: summarise driver incident reports for supervisors, about four thousand a month, low error cost, and supervisors read the originals when needed. Idea two: automatically decide compensation for damaged parcels, about three hundred a month, high error cost and disputes likely. The scorecard puts summaries first. Compensation becomes a suggest-only feature for later, once they have evidence and a review process.

### Worked example: Dubai creator agency

A worked example. A Dubai agency managing thirty creators lists candidates: caption drafts, brief summaries, sponsorship contract review, comment triage and weekly performance reports. Comment triage wins on volume, low error cost and easy checking. Report writing wins because it is weekly, slow and reviewable. Contract review is valuable but risky, so they scope it as highlight clauses for the lawyer, not approve contracts.

### FAQ: experiment everywhere or big bets?

A question from leadership teams: should we let every department experiment with AI, or pick a few big bets? Usually both, with different rules. Give everyone approved tools and simple guidelines for personal productivity, which surfaces ideas cheaply. Then run the scorecard on the most promising ideas and fund two or three focused bets with proper evaluation and owners. Broad experimentation finds opportunities; focused bets turn them into real products.

### FAQ: measuring quality value

Another question: how do we measure value if our AI feature improves quality rather than speed? Pick a quality outcome you already track: fewer errors in invoices, fewer customer re-contacts, higher conversion, fewer escalations. Measure it before and after, ideally with a comparison group. Quality improvements are often worth more than time savings, but only if you measure them in terms the business already cares about.

### Try this now

Try this now. Write down five tasks your team did this week that involved reading or writing. For each, answer two quick questions: if the AI got this wrong, how bad would it be, and how fast could someone check it? Circle the tasks where the answers are not very bad and very fast. Those circled tasks are your shortlist for the scoring workshop.

### Watch me do it

Watch me do it. I'm running a sixty-minute scoring workshop with a support team of eight. First, ten minutes of silent brainstorming: everyone writes tasks that eat their time. We get twenty-three sticky notes. We merge duplicates and end with fourteen. Next, scoring: for each task we vote one to five on the six fit dimensions, and I type the scores into the scorecard live. Two tasks rise to the top: drafting first replies to warranty emails, and tagging incoming tickets. A third, approving goodwill refunds, scores high on value but low on tolerance, so we re-scope it to suggesting a refund amount for a supervisor. Then quick sizing for the top two: volume from the ticketing system, minutes saved from an agent's honest estimate, and review time subtracted. Finally, I ask who will own discovery for each, and we leave with two named opportunities and a date to report back.

### Recap

Recap. Pick problems before models. Look for efficiency, new scale, better decisions and new experiences. Score fit on six dimensions, prioritising error tolerance and verifiability. Avoid deterministic logic and AI for its own sake. Size value net of review time. Your next step: run the scoring workshop and produce a prioritised scorecard.

## Key takeaways

- AI creates value through efficiency, new scale, better decisions and new experiences
- Score fit on language-heaviness, error tolerance, verifiability, volume, data access and measurability
- Deterministic logic, rare tasks and unresolved process problems are poor AI candidates
- Size value as volume × time saved × share handled × cost, net of review time
- Pick first features with high tolerance and easy verification

## Try it

Run a 60-minute scoring workshop with frontline colleagues on 10–20 candidate tasks and produce a prioritised opportunity scorecard.

- [Next: Automation vs augmentation: choosing the level of autonomy](https://optimizeall.com/learn/ai-product-management/automation-vs-augmentation)
- [All lessons of AI Product Management: From Idea to Reliable AI Features](https://optimizeall.com/learn/ai-product-management)
