AI Product Management: From Idea to Reliable AI FeaturesAI unit economics and pricing · Lesson 9 of 16

Unit economics: from tokens to cost per task

Article · 16 min · 9 min lecture

Video lecture

Unit economics: from tokens to cost per task

16 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 16

Unit economics for AI

  • A hit and a problem at once
  • Tokens → cost per task → margin
  • Drivers and levers

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why PMs must own AI unit economics

In traditional SaaS, the marginal cost of serving one more user is close to zero. In AI products, every request costs money (tokens, retrieval, tools, sometimes human review). A feature that delights users can quietly destroy gross margin. PMs who can translate tokens into cost per task, and cost per task into margin, make better scope, model and pricing decisions.

Tokens in one paragraph

Models process text as tokens (chunks of words). Pricing is usually per million input tokens and per million output tokens, with output typically priced higher. Token counts per word vary by language and tokeniser; for example, Arabic and Urdu text can use more tokens per word than English with some tokenisers, so measure with your provider's token counter. Images, audio and tool calls are also billed in provider-specific ways. Always use current price sheets; prices change often.

The cost-per-task formula

A "task" is the user-visible unit of value: one answered question, one drafted email, one processed invoice.

cost_per_task = Σ over model calls in the task [ input_tokens × input_price + output_tokens × output_price ]
              + retrieval/embedding costs
              + tool/API costs (search, OCR, speech)
              + human review cost × review rate
              + retries and failures overhead

Key drivers:

  • Number of calls per task. Agents may make 5–30 calls; a simple feature makes 1.
  • Context size. Long system prompts, retrieved documents and conversation history inflate input tokens on every call.
  • Output length. Verbose answers cost more and are slower.
  • Model choice. Price differences between model tiers can be an order of magnitude or more.

Levers to reduce cost (without hurting quality)

LeverHow it helpsWatch out for
Prompt cachingRepeated prompt prefixes (system prompt, shared docs) billed at a discount by many providersCache rules and discounts vary; structure prompts with stable prefixes first
Model routingCheap model for easy tasks, strong model for hard onesNeeds a reliable router and evals per route
Shorter contextRetrieve fewer, better chunks; summarise historyQuality drops if you cut the wrong context
Output limitsConcise formats, max tokensTruncated answers
Batch processingMany providers discount asynchronous batch jobsOnly for non-real-time work
Fine-tuned small modelsShorter prompts, cheaper inference at volumeUpfront and maintenance costs
Caching answersReuse answers to identical questionsStaleness, personalisation

Worked example (illustrative prices and numbers)

A support assistant answers a customer question with RAG:

  • System prompt 1,200 tokens + retrieved chunks 2,000 tokens + conversation 600 tokens + question 100 tokens = 3,900 input tokens
  • Answer: 250 output tokens
  • Assume a model priced at $1.00 per million input tokens and $5.00 per million output tokens (placeholder prices):
  • Input: 3,900 × $1.00 / 1,000,000 = $0.0039
  • Output: 250 × $5.00 / 1,000,000 = $0.00125
  • Model cost ≈ $0.0052 per answer
  • Embedding the question and vector search: small (say $0.0002).
  • 10% of answers get a human check at 1 minute of an agent costing $0.15 per minute: $0.015 average. Human review dominates cost.
  • Total ≈ $0.020 per answered question. At 300,000 questions a month ≈ $6,000.

Now optimise: prompt caching discounts the stable 1,200-token system prompt; retrieval tuned from 2,000 to 1,200 tokens; better citations cut the review rate to 5%. The new total is roughly half the original, mostly from reduced human review. Lesson: model tokens are often not the biggest cost.

Hands-on: a cost-per-task calculator

# cost_per_task.py: plug in current prices from your provider's price page (values below are placeholders)
def model_cost(inp, out, price_in_per_m, price_out_per_m, cached_inp=0, cached_price_per_m=None):
    cached_price_per_m = price_in_per_m if cached_price_per_m is None else cached_price_per_m
    return ((inp - cached_inp) * price_in_per_m + cached_inp * cached_price_per_m + out * price_out_per_m) / 1e6

def task_cost(calls, retrieval=0.0, tools=0.0, review_rate=0.0, review_minutes=0.0, cost_per_minute=0.0, retry_overhead=0.05):
    models = sum(model_cost(**c) for c in calls)
    human = review_rate * review_minutes * cost_per_minute
    return (models + retrieval + tools) * (1 + retry_overhead) + human

base = task_cost([dict(inp=3900, out=250, price_in_per_m=1.0, price_out_per_m=5.0)],
                 retrieval=0.0002, review_rate=0.10, review_minutes=1, cost_per_minute=0.15)
optimized = task_cost([dict(inp=3100, out=200, price_in_per_m=1.0, price_out_per_m=5.0, cached_inp=1200, cached_price_per_m=0.1)],
                      retrieval=0.0002, review_rate=0.05, review_minutes=1, cost_per_minute=0.15)
print(f"base ${base:.4f}/task, optimized ${optimized:.4f}/task, monthly at 300k: ${base*300_000:,.0f} -> ${optimized*300_000:,.0f}")

From cost to margin

gross margin per task = price or value attributed per task − cost per task
AI gross margin %     = (AI revenue − AI serving costs) ÷ AI revenue

Track cost per task by feature and by customer segment; heavy users can be unprofitable on flat pricing (next lesson).

Pitfalls

  • Estimating from list prices without measuring real token counts (especially for Arabic/Urdu and long conversations).
  • Ignoring agent call multiplication.
  • Forgetting human review and retries.
  • Optimising tokens while review cost dominates.

How to measure success

A cost-per-task model per AI feature, based on measured token logs and current prices, with the top three cost drivers identified and a plan for each.

Key takeaways

  • Every AI request costs money; PMs must translate tokens into cost per task and margin
  • Cost per task = model calls + retrieval + tools + human review + retries
  • Drivers: calls per task, context size, output length, model choice
  • Levers: caching, routing, shorter context, output limits, batch, fine-tuning, answer caching
  • Human review often costs more than tokens; measure real token counts per language

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. An answer uses 4,000 input tokens and 200 output tokens. Input costs $1.00/M and output $5.00/M (placeholder prices). What is the model cost?
  2. In the worked example, what was the biggest cost driver?
  3. Which lever reduces cost for a long, stable system prompt reused on every call?

Put it into practice

Build a cost-per-task model for one AI feature from real token logs and current prices; identify the top three cost drivers and one lever for each.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.