Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via APICost, scale and reliability · Lesson 12 of 19
Cost tracking, attribution and budgets
Video lecture
Cost tracking, attribution and budgets
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Cost tracking and budgets
At the end of the month, the invoice arrives and it's forty percent higher than last month. Which feature caused it? Which customer? Which change? If you can't answer in five minutes, this lesson is for you. You'll learn how to track AI costs per call, attribute them to features and customers, set budgets that actually protect you, and talk about AI spend in business terms.
0:29 Why it matters
Why does this matter? AI costs behave differently from most cloud costs. They scale with usage and prompt design, they can jump overnight when a feature goes viral or an agent loops, and prices vary hugely between models. Think of a household electricity bill. The total tells you what you spent, but a smart meter per appliance tells you it was the old freezer in the garage. Your per call ledger is that smart meter.
1:02 Usage per provider
Every provider returns usage with each response. Claude reports input and output tokens, plus cache reads and cache writes. OpenAI's Responses API reports input and output tokens, cached tokens, and reasoning tokens. Gemini reports prompt tokens, output tokens, cached tokens and thinking tokens. Here's the catch: reasoning tokens are billed as output on providers that expose them, so a thoughtful model can bill far more tokens than the visible answer suggests.
1:33 The cost ledger
So build a ledger. One row per model call, with the timestamp, provider, model, feature, tenant, a pseudonymized user id, environment, every token type, whether it was a batch request, latency, outcome and estimated cost. Calculate cost from prices loaded from configuration with an effective date, never hard coded. And once a month, reconcile your estimates against the provider's usage and billing reports, so you know your numbers are right.
2:03 Simple example: three features
A simple example. A solo developer runs a writing assistant with three features: rewrite, summarize and translate. After adding a ledger, she sees translation is a third of her requests but two thirds of her cost, because users paste whole documents. She adds an input size limit with a friendly message and moves long documents to a cheaper model. Her bill drops, and nobody complains.
2:31 Hands-on: cost meter
The lesson's code gives you a provider agnostic cost meter. Small adapter functions convert each provider's usage object into one common shape: input, cached input, cache writes and output. A cost function applies prices from your config, with a batch multiplier where relevant. And a record function writes a row with feature and tenant tags. Check each provider's token accounting, like whether cached tokens are included in the input total, and adjust the adapters if needed.
3:04 Budgets that protect
Now budgets that protect you. Per request caps, like max tokens, input size limits and agent step caps. Per user and per tenant quotas enforced in your backend. Per feature budgets with alerts at fifty, eighty and a hundred percent. Provider side spend limits and separate projects. And anomaly alerts when cost per request or per tenant jumps compared with the recent average. One customer's runaway loop shouldn't become everyone's problem.
3:35 Per-tenant limits
Here's a practical detail that saves real money: per tenant limits. If you sell an AI feature to many customers, one customer's integration can accidentally call you in a loop overnight. Without limits, their mistake becomes your invoice. Give every tenant a daily token or cost allowance, enforced in your backend before calling the provider. When they hit it, return a clear message and alert your team. Most customers will never notice the limit; the one who hits it will thank you for catching their bug.
4:12 Business example: agency billing
Now a realistic business example, with illustrative numbers. An agency working across the UK and Pakistan runs AI workflows for twenty five clients. With the ledger tagged by client and feature, finance produces a monthly AI cost report per client, adds a margin, and invoices it. The report also revealed that one client's weekly report feature cost about ten times more than similar clients, because it attached a sixty page PDF every run. Caching and page selection fixed it.
4:46 Unit economics
Finally, speak the language of the business: unit economics. Translate tokens into cost per support conversation resolved, per lead enriched, per product description published, or per active user per month. Then compare that with the value: time saved, revenue, conversion. This is how you decide where a bigger model is worth it, where caching or batch deserve engineering time, and how to price AI features for your own customers.
5:16 Common mistakes
Common mistakes. Relying only on the provider invoice. Hard coded prices that silently go stale. Ignoring reasoning tokens and cache writes. And no per tenant limits, so one customer's loop becomes everyone's bill. Each of these is cheap to prevent today and painful to untangle after a surprising invoice arrives.
5:38 Showing costs
A common question: should you show AI costs to your customers or users? For internal teams, yes: showing each team its monthly AI spend changes behavior quickly. For external customers, it depends on your pricing model. Many products bundle AI into plan tiers with fair use limits, while agencies often pass through costs with a margin. Either way, your ledger gives you the numbers to decide pricing with confidence rather than guesswork.
6:09 Forecast monthly
One last habit: forecast. Take the last four weeks of ledger data per feature, work out cost per outcome and outcomes per week, and project forward with known changes, like a new client, a marketing campaign or a model switch. Add a buffer for growth and price changes, share it with finance, and compare forecast with actual every month. The gap tells you which assumption was wrong, and over time your forecasts become boringly accurate, which is exactly what finance wants.
6:44 Deeper: agency AI billing (illustrative)
Let's deepen the UK and Pakistan agency's billing example with illustrative numbers. Twenty five clients, three AI features each. The ledger showed one client's weekly report costing about ten times the median, caused by a sixty page PDF attached to every run. After caching the document and selecting only the relevant pages, that client's cost dropped in line with the others. Finance then introduced a simple pricing rule: AI usage passed through at cost plus a fixed margin, shown as a line on each monthly invoice. Clients appreciated the transparency, and the agency stopped absorbing AI costs it couldn't see before.
7:28 Watch me do it: provider-agnostic cost meter
Watch me do it. Let's walk through the cost meter. Prices load from a JSON file with an effective date, never hard coded. The usage data class has provider, model, input, cached input, cache writes and output. Three adapter functions convert each provider's usage object. From Claude reads input tokens, cache read and cache creation tokens and output tokens. From OpenAI subtracts cached tokens from input so they aren't double counted. From Gemini subtracts cached tokens from the prompt count and adds thinking tokens to output, because they're billed as output. The cost function multiplies each part by its price per million, with a batch multiplier when relevant. Record builds a row with the feature, tenant, batch flag, cost and timestamp, and prints it, standing in for your data warehouse. I record one call from each provider for the feature product copy and tenant noor, and three comparable rows appear.
8:33 Recap + try this now
Quick recap. Invoices show totals, so build your own per call ledger with feature and tenant tags, prices from config, and monthly reconciliation. Enforce budgets at every level, and express spend as cost per business outcome. Try this now: add a cost ledger to one application, set one budget alert, and calculate the cost per outcome for one feature, like cost per resolved support conversation.
Why cost tracking needs design
AI spend behaves differently from most cloud costs: it scales with usage and prompt design, can jump overnight when a feature goes viral or a loop misbehaves, and varies by model by an order of magnitude or more. Provider invoices tell you what you spent, not which feature, customer or change caused it. You need your own attribution.
The usage data you get per request
- Claude:
usage.input_tokens,output_tokens,cache_creation_input_tokens,cache_read_input_tokens(plus server-tool usage where applicable). - OpenAI Responses:
usage.input_tokens,output_tokens,input_tokens_details.cached_tokens,output_tokens_details.reasoning_tokens. - Gemini:
usage_metadata.prompt_token_count,candidates_token_count,cached_content_token_count,thoughts_token_count,total_token_count.
Reasoning ("thinking") tokens are billed as output on providers that expose them; long-thinking models can produce many more billed tokens than the visible answer suggests.
A cost ledger
Record one row per model call:
timestamp, request_id, provider, model, feature, tenant_id, user_id (pseudonymized), environment,
input_tokens, cached_input_tokens, cache_write_tokens, output_tokens, reasoning_tokens,
batch (bool), estimated_cost, latency_ms, outcome (success/error/refusal)Estimated cost = tokens × prices loaded from configuration (with an effective date), not hard-coded. Reconcile monthly against provider invoices and usage/cost reports (Anthropic's usage and cost Admin API, OpenAI's usage dashboards and APIs, Google Cloud billing export).
Hands-on: a provider-agnostic cost meter
import json, time
from dataclasses import dataclass, asdict
# Prices per 1M tokens, loaded from config with an effective date. Values here are placeholders.
PRICES = json.load(open("prices.json")) # {"claude-sonnet-5": {"in": ..., "out": ..., "cache_read": ..., "cache_write": ...}, ...}
@dataclass
class Usage:
provider: str
model: str
input_tokens: int = 0
cached_input_tokens: int = 0
cache_write_tokens: int = 0
output_tokens: int = 0
def from_claude(model, u) -> Usage:
return Usage("anthropic", model, u.input_tokens, u.cache_read_input_tokens or 0,
u.cache_creation_input_tokens or 0, u.output_tokens)
def from_openai(model, u) -> Usage:
cached = getattr(getattr(u, "input_tokens_details", None), "cached_tokens", 0) or 0
return Usage("openai", model, u.input_tokens - cached, cached, 0, u.output_tokens)
def from_gemini(model, m) -> Usage:
cached = m.cached_content_token_count or 0
out = (m.candidates_token_count or 0) + (m.thoughts_token_count or 0)
return Usage("google", model, (m.prompt_token_count or 0) - cached, cached, 0, out)
def cost(u: Usage, batch: bool = False) -> float:
p = PRICES[u.model]
c = (u.input_tokens * p["in"] + u.cached_input_tokens * p.get("cache_read", p["in"])
+ u.cache_write_tokens * p.get("cache_write", p["in"]) + u.output_tokens * p["out"]) / 1_000_000
return c * (p.get("batch_multiplier", 0.5) if batch else 1.0)
def record(u: Usage, feature: str, tenant: str, batch: bool = False):
row = {**asdict(u), "feature": feature, "tenant": tenant, "batch": batch,
"cost": round(cost(u, batch), 6), "ts": time.time()}
print(json.dumps(row)) # send to your warehouse / observability pipeline insteadCheck each provider's token-accounting semantics (for example, whether cached tokens are included in the input total) against current docs, and adjust the adapters.
Budgets and guardrails
- Per-request caps:
max_tokens/max_output_tokens, input size limits, step caps for agents. - Per-user and per-tenant quotas: daily token or cost limits enforced in your backend.
- Per-feature budgets with alerts at 50%, 80% and 100%.
- Provider-side limits: spend limits and alerts in each console, separate projects per product.
- Anomaly alerts: cost per request or per tenant jumping versus the trailing average.
Unit economics
Translate tokens into business terms: cost per support conversation resolved, per lead enriched, per product description published, per active user per month. Compare with the value created (time saved, revenue, conversion). This is how you decide where to use bigger models, where caching or batch are worth engineering time, and how to price AI features.
Worked example: an agency billing clients for AI usage
A UK-and-Pakistan agency runs AI workflows for 25 clients. With the ledger tagged by tenant_id and feature, finance produces a monthly per-client AI cost report, adds a margin, and invoices it. They also discovered that one client's "weekly report" feature cost ten times more than similar clients (illustrative) because it attached a 60-page PDF every run; caching and page selection fixed it.
Forecasting next month
A simple forecast is enough for most teams: take the last four weeks of ledger data per feature, compute cost per outcome and outcomes per week, and project forward with known changes (a new client onboarding, a marketing campaign, a model switch). Add a buffer for growth and for price changes. Share the forecast with finance alongside the budget alerts, and compare forecast to actual each month; the gap tells you which assumptions to fix.
Pitfalls
- Relying only on the provider invoice.
- Hard-coded prices that silently go stale.
- Ignoring reasoning tokens and cache writes.
- No per-tenant limits, so one customer's loop becomes everyone's bill.
Measuring success
Share of spend attributed to a feature and tenant (target near 100%), forecast accuracy, number of budget alerts acted on, and cost per business outcome over time.
Key takeaways
- Provider invoices show totals; you need your own attribution by feature, tenant and change.
- Record usage per call, including cached, cache-write and reasoning tokens, with prices from configuration.
- Reconcile estimates monthly with provider usage and billing reports.
- Enforce budgets per request, user, tenant and feature, with anomaly alerts.
- Express AI cost as unit economics tied to business outcomes.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Add a per-call cost ledger with feature and tenant tags to one application, set one budget alert, and produce a cost-per-outcome figure for one feature.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.