AI Product Management: From Idea to Reliable AI FeaturesAI unit economics and pricing · Lesson 9 of 16
Unit economics: from tokens to cost per task
Video lecture
Unit economics: from tokens to cost per task
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Unit economics for AI
Here is a story that plays out in many AI startups. A feature launches, users love it, usage triples, and then finance calls a meeting. Every answer costs money, the heaviest users cost more than they pay, and gross margin is sinking. The feature was a hit and a problem at the same time. In this lesson you will learn to turn tokens into cost per task, find the real cost drivers, and pull the levers that cut cost without hurting quality.
0:36 Why this is new
Why is this new for PMs? In traditional software, serving one more user costs almost nothing. In AI products, every request costs money: tokens, retrieval, tools and sometimes a person checking the output. Here is an analogy. Traditional software is like a streaming service: one more viewer barely costs anything. AI features are more like a restaurant: every plate has ingredient costs. You would never price a menu without knowing what each dish costs to make.
1:09 Tokens in brief
A quick primer on tokens. Models read and write text in chunks called tokens. Providers charge per million input tokens and per million output tokens, and output usually costs more. Token counts per word vary by language and tokeniser, and Arabic or Urdu text can use more tokens per word than English with some tokenisers, so measure with your provider's counter rather than guessing. And prices change often, so always use the current price page.
1:42 Cost per task
Now the formula. A task is the unit users care about: one answered question, one drafted email, one processed invoice. Cost per task is the sum over every model call in that task, input tokens times the input price plus output tokens times the output price, plus retrieval and embedding costs, plus tool costs like search, OCR or speech, plus human review cost times the review rate, plus overhead for retries and failures.
2:14 Four cost drivers
Four drivers dominate. Calls per task: a simple feature makes one call, an agent may make five to thirty. Context size: long system prompts, retrieved documents and conversation history inflate input on every call. Output length: verbose answers cost more and feel slower. And model choice: price differences between model tiers can be an order of magnitude or more. When cost surprises you, it is almost always one of these four.
2:45 Simple example (placeholder prices)
A simple example with placeholder prices. One answer uses four thousand input tokens and two hundred output tokens. Input costs one dollar per million tokens, output five dollars per million. Input: four thousand times one, divided by a million, is zero point four cents. Output: two hundred times five, divided by a million, is zero point one cent. Total: half a cent per answer. Multiply by your monthly volume, and you have your model bill for that feature.
3:19 Realistic example (illustrative)
Now a realistic example, with illustrative numbers. A support assistant with retrieval uses about thirty-nine hundred input tokens and two hundred and fifty output tokens per answer: about half a cent in model cost. But ten per cent of answers get a one-minute human check by an agent costing fifteen cents a minute. That adds one and a half cents on average, three times the model cost. Total: about two cents per answer, around six thousand dollars a month at three hundred thousand questions.
3:55 Optimisation (illustrative)
Then they optimise. They move the stable system prompt to the front so prompt caching discounts it. They tune retrieval to send fewer, better chunks. And better citations let reviewers check less often, cutting the review rate to five per cent. The new cost is roughly half the original, and most of the saving came from less human review, not fewer tokens. The lesson: find the real driver before you optimise.
4:25 Cost levers
The full set of levers: prompt caching for stable prefixes; routing easy tasks to cheaper models; shorter context through better retrieval and summarised history; output limits and concise formats; batch processing for work that is not real time; fine-tuned small models at high volume; and caching answers to identical questions, if freshness allows. Each has a catch, and the lesson's table lists them. The calculator script lets you model each lever before you build it.
4:58 Common mistakes
Common mistakes. Estimating from list prices without measuring real token counts, especially for Arabic, Urdu and long conversations. Forgetting that agents multiply calls. Forgetting human review and retries. And optimising tokens when review cost dominates. Finally, track cost per task by feature and by customer segment, because averages hide heavy users who may be unprofitable.
5:22 Another example: travel booking agent (illustrative)
Another example, with illustrative numbers: an AI agent. A travel company's booking agent makes, on average, twelve model calls per booking conversation: understanding the request, searching, comparing options, checking policies, drafting and confirming. Each call carries a growing conversation history. The team discovers that by the last call, input tokens are five times larger than the first. Summarising history after every few steps and using a cheaper model for search steps cuts cost per booking substantially, without hurting completion rates.
5:57 Keep it current
How often should you revisit cost per task? Monthly at minimum, and whenever a provider changes prices, you change models, or usage patterns shift. Put cost per task on the same dashboard as quality, with an owner. The best AI teams treat cost as a product metric, reviewed alongside quality, not as a finance surprise at the end of the quarter.
6:24 FAQ: show AI costs to customers?
A question PMs ask: should we show AI costs to customers? Usually not raw token costs, which mean little to most users. Show usage in units they understand, like documents processed, answers generated or credits remaining. Internally, though, make cost per task visible to the whole product team. When designers and engineers can see what a feature costs per use, they make smarter choices about context, output length and model selection.
6:55 Try this now
Try this now. Pull a week of token logs for one AI feature. Compute average input tokens, output tokens and calls per task. Multiply by today's prices from your provider's price page, add an estimate for human review, and write the cost per task on a sticky note on your monitor. Then circle the biggest term. That is where your next optimisation belongs.
7:22 Watch me do it
Watch me do it. I export one week of logs for our AI support assistant and compute averages per answered question: two calls, four thousand two hundred input tokens in total, three hundred output tokens. I take today's prices from the provider page and note the date. Model cost per answer: about half a cent, with these placeholder prices. Retrieval adds a little. Then the part people forget: I check the review queue; twelve per cent of answers get a human check averaging ninety seconds. At our loaded agent cost, that is more than the model cost. Now I test levers in the calculator. Moving the stable system prompt to the front for caching saves a bit. Cutting retrieved chunks from six to four saves more. Adding claim-level citations, which reviewers say would halve their checking time, saves the most. I write the top three drivers and the plan on one slide.
8:28 Recap
Recap. Every AI request costs money, so PMs own cost per task. Add up model calls, retrieval, tools, review and retries. Watch calls, context, output length and model tier. Pull the right lever for the biggest driver, which is often human review. And track cost by feature and segment. Next, we turn cost into pricing.
Why PMs must own AI unit economics
In traditional SaaS, the marginal cost of serving one more user is close to zero. In AI products, every request costs money (tokens, retrieval, tools, sometimes human review). A feature that delights users can quietly destroy gross margin. PMs who can translate tokens into cost per task, and cost per task into margin, make better scope, model and pricing decisions.
Tokens in one paragraph
Models process text as tokens (chunks of words). Pricing is usually per million input tokens and per million output tokens, with output typically priced higher. Token counts per word vary by language and tokeniser; for example, Arabic and Urdu text can use more tokens per word than English with some tokenisers, so measure with your provider's token counter. Images, audio and tool calls are also billed in provider-specific ways. Always use current price sheets; prices change often.
The cost-per-task formula
A "task" is the user-visible unit of value: one answered question, one drafted email, one processed invoice.
cost_per_task = Σ over model calls in the task [ input_tokens × input_price + output_tokens × output_price ]
+ retrieval/embedding costs
+ tool/API costs (search, OCR, speech)
+ human review cost × review rate
+ retries and failures overheadKey drivers:
- Number of calls per task. Agents may make 5–30 calls; a simple feature makes 1.
- Context size. Long system prompts, retrieved documents and conversation history inflate input tokens on every call.
- Output length. Verbose answers cost more and are slower.
- Model choice. Price differences between model tiers can be an order of magnitude or more.
Levers to reduce cost (without hurting quality)
| Lever | How it helps | Watch out for |
|---|---|---|
| Prompt caching | Repeated prompt prefixes (system prompt, shared docs) billed at a discount by many providers | Cache rules and discounts vary; structure prompts with stable prefixes first |
| Model routing | Cheap model for easy tasks, strong model for hard ones | Needs a reliable router and evals per route |
| Shorter context | Retrieve fewer, better chunks; summarise history | Quality drops if you cut the wrong context |
| Output limits | Concise formats, max tokens | Truncated answers |
| Batch processing | Many providers discount asynchronous batch jobs | Only for non-real-time work |
| Fine-tuned small models | Shorter prompts, cheaper inference at volume | Upfront and maintenance costs |
| Caching answers | Reuse answers to identical questions | Staleness, personalisation |
Worked example (illustrative prices and numbers)
A support assistant answers a customer question with RAG:
- System prompt 1,200 tokens + retrieved chunks 2,000 tokens + conversation 600 tokens + question 100 tokens = 3,900 input tokens
- Answer: 250 output tokens
- Assume a model priced at $1.00 per million input tokens and $5.00 per million output tokens (placeholder prices):
- Input: 3,900 × $1.00 / 1,000,000 = $0.0039
- Output: 250 × $5.00 / 1,000,000 = $0.00125
- Model cost ≈ $0.0052 per answer
- Embedding the question and vector search: small (say $0.0002).
- 10% of answers get a human check at 1 minute of an agent costing $0.15 per minute: $0.015 average. Human review dominates cost.
- Total ≈ $0.020 per answered question. At 300,000 questions a month ≈ $6,000.
Now optimise: prompt caching discounts the stable 1,200-token system prompt; retrieval tuned from 2,000 to 1,200 tokens; better citations cut the review rate to 5%. The new total is roughly half the original, mostly from reduced human review. Lesson: model tokens are often not the biggest cost.
Hands-on: a cost-per-task calculator
# cost_per_task.py: plug in current prices from your provider's price page (values below are placeholders)
def model_cost(inp, out, price_in_per_m, price_out_per_m, cached_inp=0, cached_price_per_m=None):
cached_price_per_m = price_in_per_m if cached_price_per_m is None else cached_price_per_m
return ((inp - cached_inp) * price_in_per_m + cached_inp * cached_price_per_m + out * price_out_per_m) / 1e6
def task_cost(calls, retrieval=0.0, tools=0.0, review_rate=0.0, review_minutes=0.0, cost_per_minute=0.0, retry_overhead=0.05):
models = sum(model_cost(**c) for c in calls)
human = review_rate * review_minutes * cost_per_minute
return (models + retrieval + tools) * (1 + retry_overhead) + human
base = task_cost([dict(inp=3900, out=250, price_in_per_m=1.0, price_out_per_m=5.0)],
retrieval=0.0002, review_rate=0.10, review_minutes=1, cost_per_minute=0.15)
optimized = task_cost([dict(inp=3100, out=200, price_in_per_m=1.0, price_out_per_m=5.0, cached_inp=1200, cached_price_per_m=0.1)],
retrieval=0.0002, review_rate=0.05, review_minutes=1, cost_per_minute=0.15)
print(f"base ${base:.4f}/task, optimized ${optimized:.4f}/task, monthly at 300k: ${base*300_000:,.0f} -> ${optimized*300_000:,.0f}")From cost to margin
gross margin per task = price or value attributed per task − cost per task
AI gross margin % = (AI revenue − AI serving costs) ÷ AI revenueTrack cost per task by feature and by customer segment; heavy users can be unprofitable on flat pricing (next lesson).
Pitfalls
- Estimating from list prices without measuring real token counts (especially for Arabic/Urdu and long conversations).
- Ignoring agent call multiplication.
- Forgetting human review and retries.
- Optimising tokens while review cost dominates.
How to measure success
A cost-per-task model per AI feature, based on measured token logs and current prices, with the top three cost drivers identified and a plan for each.
Key takeaways
- Every AI request costs money; PMs must translate tokens into cost per task and margin
- Cost per task = model calls + retrieval + tools + human review + retries
- Drivers: calls per task, context size, output length, model choice
- Levers: caching, routing, shorter context, output limits, batch, fine-tuning, answer caching
- Human review often costs more than tokens; measure real token counts per language
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Build a cost-per-task model for one AI feature from real token logs and current prices; identify the top three cost drivers and one lever for each.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.