Skip to content

Building AI Products & Workflows · Unit economics and UX patterns for AI features · lesson 13 of 18 · 11 min

Unit economics of AI features

Why AI changes product economics

Traditional software features have near-zero marginal cost per use. AI features have variable costs per interaction: tokens, model calls, retrieval, tool calls and sometimes human review. A feature that delights users can lose money at scale if nobody models the economics. Prices change frequently and vary by provider, so build models that you can update, rather than relying on fixed numbers.

Cost components

  • Model usage: input tokens, output tokens, reasoning tokens, images or audio processed.
  • Retrieval and storage: embeddings, vector database, document processing.
  • Tool and API calls: third-party services the feature uses.
  • Infrastructure: hosting, observability, evaluation runs.
  • Human costs: review, escalation handling, quality sampling.
  • Development and maintenance: engineering, prompt and evaluation upkeep.

A simple cost model (illustrative)

Cost per task = model cost + retrieval cost + tool cost + human review cost

model cost    = calls per task x (input tokens x input price + output tokens x output price)
review cost   = review rate x minutes per review x loaded cost per minute

Monthly cost  = cost per task x tasks per month + fixed platform costs
Value         = time saved x loaded cost + revenue impact + quality gains

Fill it with your own measured numbers from prototypes and pilots. Human review is often the largest cost component, which is why improving quality (reducing review needs) can matter more than cheaper tokens.

Cost per outcome

Measure cost per successful outcome: cost per resolved ticket, per qualified lead, per processed invoice. A cheaper model that fails more often may cost more per outcome after retries, escalations and human time.

Levers to improve unit economics

  • Right-size models: small, fast models for simple steps; larger ones only where needed.
  • Route by difficulty: easy cases automated cheaply, hard ones escalated.
  • Prompt caching: stable prefixes (instructions, documents) cached where providers support it.
  • Trim context: shorter prompts, fewer retrieved passages, concise tool results.
  • Batch processing: for non-urgent work, batch endpoints are often cheaper.
  • Reduce retries: better prompts and structured outputs reduce validation failures.
  • Improve quality to reduce review: often the biggest lever.

Pricing AI features

If you sell a product with AI features, pricing must cover variable costs:

  • Bundled: included in existing plans; watch heavy users.
  • Usage-based: credits or per-use fees; transparent but can deter use.
  • Tiered: AI features in higher tiers; usage limits per tier.
  • Outcome-based: pay per resolution or result; aligns value but shifts risk to you.

Protect margins with fair-use limits, rate limits and monitoring of heavy users, and communicate limits clearly to customers.

Worked example: invoice processing

An operations team models invoice extraction:

  • Model and processing cost per invoice: small.
  • Human review: 20% of invoices reviewed at a few minutes each, which dominates cost.
  • Improving field accuracy reduces the review rate to 10% for trusted suppliers.
  • Result: total cost per invoice falls substantially more from the review reduction than from switching to a cheaper model.

The team invests in evaluation and quality improvements rather than chasing the lowest token price.

Budget governance

  • Set budgets and alerts per feature.
  • Track cost per outcome monthly, alongside value delivered.
  • Review model choices when providers change prices or release new models.
  • Include evaluation and monitoring costs in the business case; they are not optional.

Hands-on: a cost-per-outcome model you can update in minutes

Keep prices in a config file you update when providers change them (check each provider's pricing page), and keep volumes, token counts and review rates from your own measurements. All numbers below are illustrative placeholders.

# unit_economics.py
PRICES = {  # per million tokens, illustrative only: update from your provider's pricing page
    "large": {"input": 5.00, "output": 25.00, "cache_read_multiplier": 0.1},
    "small": {"input": 1.00, "output": 5.00, "cache_read_multiplier": 0.1},
}
BATCH_DISCOUNT = 0.5          # many providers discount asynchronous batch processing; verify yours

def model_cost(tier, in_tok, out_tok, cached_share=0.0, batch=False, calls=1):
    p = PRICES[tier]
    effective_in = in_tok * ((1 - cached_share) + cached_share * p["cache_read_multiplier"])
    cost = calls * (effective_in * p["input"] + out_tok * p["output"]) / 1_000_000
    return cost * (BATCH_DISCOUNT if batch else 1)

def cost_per_outcome(s):
    ai = model_cost(s["tier"], s["in_tok"], s["out_tok"], s["cached_share"], s["batch"], s["calls"])
    review = s["review_rate"] * s["review_minutes"] * s["loaded_cost_per_min"]
    retries = ai * s["retry_rate"]
    per_task = ai + retries + review + s["other_per_task"]
    return per_task / s["success_rate"]                     # failures still cost money

scenarios = {
    "baseline":        dict(tier="large", in_tok=6000, out_tok=800, cached_share=0.0, batch=False, calls=1,
                            review_rate=0.20, review_minutes=4, loaded_cost_per_min=0.60, retry_rate=0.05,
                            other_per_task=0.002, success_rate=0.92),
    "cache+route":     dict(tier="small", in_tok=6000, out_tok=800, cached_share=0.7, batch=False, calls=1,
                            review_rate=0.20, review_minutes=4, loaded_cost_per_min=0.60, retry_rate=0.08,
                            other_per_task=0.002, success_rate=0.88),
    "quality-first":   dict(tier="large", in_tok=6000, out_tok=800, cached_share=0.7, batch=True, calls=1,
                            review_rate=0.10, review_minutes=4, loaded_cost_per_min=0.60, retry_rate=0.03,
                            other_per_task=0.002, success_rate=0.96),
}
for name, s in scenarios.items():
    print(f"{name:14} cost per successful outcome: {cost_per_outcome(s):.3f}")

Run it and notice which lever moves the result most. In many real features it is the review rate, not the token price. Then multiply by monthly volume for low, expected and high adoption, and compare with the value side (time saved times loaded cost, revenue impact). Update the inputs monthly from your monitoring data; the model is only as good as its measurements.

Going further

Build a simple unit-economics model in a spreadsheet with inputs (volume, tokens, prices, review rate, time saved) and scenarios (low, expected, high adoption). Update it with real data after launch. It turns debates about "is AI worth it?" into specific, testable assumptions.

Video lecture: Unit economics of AI features

Lecture coming soon · 15 chapters · about 8 minutes. Read the full transcript below.

  1. Unit economics of AI features
  2. Analogy: bridge vs taxi
  3. Cost components
  4. Cost per successful outcome
  5. Levers
  6. Simple example (illustrative)
  7. Worked example: invoice extraction
  8. Business example (illustrative)
  9. Pricing and budgets
  10. Hands-on in the lesson
  11. Common mistakes
  12. How you'll know it's healthy
  13. Watch me do it: cost model
  14. Recap
  15. Try this now (30 minutes)

Lecture transcript

Unit economics of AI features

Traditional software features cost almost nothing each time someone clicks them. AI features are different. Every interaction spends tokens, calls models and tools, and sometimes needs a person to review the result. A feature your users love can quietly lose money at scale if nobody models the economics. In this lesson you'll learn the cost components, how to calculate cost per successful outcome, the levers that actually move it, and how to price AI features you sell.

Analogy: bridge vs taxi

Here's an analogy. Traditional software is like a toll-free bridge you've already built: more cars cost you almost nothing. AI features are more like a taxi service: every trip burns fuel, and some trips need a second driver to check the route. You can make taxis profitable, but only if you know the cost per trip and price accordingly.

Cost components

Here are the components. Model usage: input tokens, output tokens, reasoning tokens, and images or audio processed. Retrieval and storage: embeddings, vector databases and document processing. Tool and API calls to third-party services. Infrastructure: hosting, observability and evaluation runs. Human costs: review, escalations and quality sampling. And development and maintenance: engineering time, prompt and evaluation upkeep. Prices change often and vary by provider, so build a model you can update, rather than memorising numbers.

Cost per successful outcome

A simple model works. Cost per task equals model cost, plus retrieval, plus tools, plus human review. Model cost is calls per task times tokens times price. Review cost is review rate times minutes per review times loaded cost per minute. But the number that matters is cost per successful outcome: per resolved ticket, per qualified lead, per processed invoice. Divide by your success rate, because failures still cost money. A cheaper model that fails more often can cost more per outcome after retries, escalations and human time.

Levers

Now the levers. Right-size models: small, fast ones for simple steps, larger ones only where they're needed. Route by difficulty. Use prompt caching for stable prefixes like instructions and documents, which providers bill at a fraction of the normal input price. Trim context: shorter prompts, fewer passages, concise tool results. Use batch processing for work that isn't urgent; many providers discount it. Reduce retries with structured outputs. And improve quality to reduce human review, which is often the biggest lever of all.

Simple example (illustrative)

A simple example with illustrative numbers. An AI step costs about two cents per email in model usage. One in ten emails needs a person to spend three minutes fixing it, at about fifty cents a minute. That's fifteen cents of review spread across every email, seven times the model cost. Improve quality so only one in twenty needs fixing, and you save more than switching to a model half the price.

Worked example: invoice extraction

Here's a worked example. An operations team models invoice extraction. Model and processing cost per invoice is small. Human review of one in five invoices, a few minutes each, dominates the cost. By improving field accuracy, they reduce review to one in ten for trusted suppliers. Total cost per invoice falls far more from that change than it would have from switching to a cheaper model. So they invest in evaluation and quality, not in chasing the lowest token price.

Business example (illustrative)

More on the invoice team, illustrative. They process about four thousand invoices a month. Model cost is a few pence each; review at twenty percent, four minutes each, is the dominant cost. Better field accuracy cut review to ten percent for trusted suppliers, saving roughly fifty review hours a month. Switching to a model half the price would have saved far less, and their evaluation showed it would have raised the review rate.

Pricing and budgets

If you sell a product with AI features, pricing must cover variable costs. Bundling into existing plans is simple, but watch heavy users. Usage-based credits are transparent but can deter use. Tiered plans put AI features in higher tiers with limits. Outcome-based pricing, per resolution or per result, aligns with value but shifts risk to you. Protect margins with fair-use limits, rate limits and monitoring of heavy users, and communicate limits clearly. And govern budgets: per-feature budgets and alerts, monthly cost per outcome alongside value delivered, and a review whenever providers change prices or launch new models.

Hands-on in the lesson

The hands-on section gives you a small Python cost model with prices in a config you update from your provider's pricing page, and parameters for caching share, batch discounts, retries, review rate and success rate. It prints cost per successful outcome for three illustrative scenarios, so you can see which lever moves the result most. In many real features, that's the review rate. Then multiply by volume for low, expected and high adoption, and compare with value.

Common mistakes

Common mistakes. Modelling only token costs. Using list prices without checking caching or batch discounts. Measuring cost per call instead of per successful outcome. Ignoring heavy users who drive most of the cost. Forgetting evaluation and monitoring in the business case. And never updating the model after launch, while prices, volumes and review rates all change.

How you'll know it's healthy

How will you know your economics are healthy? Cost per successful outcome is tracked monthly and trending down or stable. It's comfortably below the value delivered per outcome. Heavy users are visible and within fair-use limits. And your pricing, if you sell the feature, covers variable costs at the high-adoption scenario, not just the expected one.

Watch me do it: cost model

Watch me do it. I open unit economics dot py. First, prices live in a config dictionary per model tier, with a cache read multiplier, and a batch discount constant. I update these from the provider's pricing page. Next, model cost applies the cached share of input at the reduced rate, multiplies by calls, and applies the batch discount if used. Then cost per outcome adds model cost, retries, review cost from review rate times minutes times cost per minute, and other costs, and divides by the success rate, because failed tasks still cost money. I run three scenarios. Baseline is the most expensive per successful outcome. Cache plus route is cheaper per call but its lower success rate eats some of the saving. Quality first, with caching, batch and half the review rate, comes out lowest. I multiply by monthly volume for three adoption levels and paste the results into the business case.

Recap

To recap: AI features have variable per-use costs. Model all the components, including human review and maintenance. Measure cost per successful outcome, pull the levers that matter, often quality before token price, and price AI features to cover variable costs with sensible limits. Your next step is to build a unit-economics model for one feature with low, expected and high adoption, including review costs. Next: UX patterns that build calibrated trust.

Try this now (30 minutes)

Try this now. Take one AI feature and fill in the cost model from the lesson with your own numbers: tokens per task, price, caching share, review rate, minutes per review, loaded cost and success rate. Calculate cost per successful outcome. Then change one lever at a time, model tier, caching or review rate, and note which one moves the number most. That's where to focus.

Key takeaways

  • AI features have variable per-use costs; model them before scaling.
  • Costs include model usage, retrieval, tools, infrastructure, human review and maintenance.
  • Measure cost per successful outcome; human review is often the biggest cost lever.
  • Improve economics with right-sizing, routing, caching, trimming, batching and quality; price AI features to cover variable costs.

Try it

Build a spreadsheet unit-economics model for one AI feature with low, expected and high adoption scenarios, including human review costs.