Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via APICost, scale and reliability · Lesson 12 of 19

Cost tracking, attribution and budgets

Article · 14 min · 9 min lecture

Video lecture

Cost tracking, attribution and budgets

16 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 16

Cost tracking and budgets

  • Why invoices aren't enough
  • A per-call ledger
  • Budgets and alerts
  • Unit economics

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why cost tracking needs design

AI spend behaves differently from most cloud costs: it scales with usage and prompt design, can jump overnight when a feature goes viral or a loop misbehaves, and varies by model by an order of magnitude or more. Provider invoices tell you what you spent, not which feature, customer or change caused it. You need your own attribution.

The usage data you get per request

  • Claude: usage.input_tokens, output_tokens, cache_creation_input_tokens, cache_read_input_tokens (plus server-tool usage where applicable).
  • OpenAI Responses: usage.input_tokens, output_tokens, input_tokens_details.cached_tokens, output_tokens_details.reasoning_tokens.
  • Gemini: usage_metadata.prompt_token_count, candidates_token_count, cached_content_token_count, thoughts_token_count, total_token_count.

Reasoning ("thinking") tokens are billed as output on providers that expose them; long-thinking models can produce many more billed tokens than the visible answer suggests.

A cost ledger

Record one row per model call:

timestamp, request_id, provider, model, feature, tenant_id, user_id (pseudonymized), environment,
input_tokens, cached_input_tokens, cache_write_tokens, output_tokens, reasoning_tokens,
batch (bool), estimated_cost, latency_ms, outcome (success/error/refusal)

Estimated cost = tokens × prices loaded from configuration (with an effective date), not hard-coded. Reconcile monthly against provider invoices and usage/cost reports (Anthropic's usage and cost Admin API, OpenAI's usage dashboards and APIs, Google Cloud billing export).

Hands-on: a provider-agnostic cost meter

import json, time
from dataclasses import dataclass, asdict

# Prices per 1M tokens, loaded from config with an effective date. Values here are placeholders.
PRICES = json.load(open("prices.json"))   # {"claude-sonnet-5": {"in": ..., "out": ..., "cache_read": ..., "cache_write": ...}, ...}

@dataclass
class Usage:
    provider: str
    model: str
    input_tokens: int = 0
    cached_input_tokens: int = 0
    cache_write_tokens: int = 0
    output_tokens: int = 0

def from_claude(model, u) -> Usage:
    return Usage("anthropic", model, u.input_tokens, u.cache_read_input_tokens or 0,
                 u.cache_creation_input_tokens or 0, u.output_tokens)

def from_openai(model, u) -> Usage:
    cached = getattr(getattr(u, "input_tokens_details", None), "cached_tokens", 0) or 0
    return Usage("openai", model, u.input_tokens - cached, cached, 0, u.output_tokens)

def from_gemini(model, m) -> Usage:
    cached = m.cached_content_token_count or 0
    out = (m.candidates_token_count or 0) + (m.thoughts_token_count or 0)
    return Usage("google", model, (m.prompt_token_count or 0) - cached, cached, 0, out)

def cost(u: Usage, batch: bool = False) -> float:
    p = PRICES[u.model]
    c = (u.input_tokens * p["in"] + u.cached_input_tokens * p.get("cache_read", p["in"])
         + u.cache_write_tokens * p.get("cache_write", p["in"]) + u.output_tokens * p["out"]) / 1_000_000
    return c * (p.get("batch_multiplier", 0.5) if batch else 1.0)

def record(u: Usage, feature: str, tenant: str, batch: bool = False):
    row = {**asdict(u), "feature": feature, "tenant": tenant, "batch": batch,
           "cost": round(cost(u, batch), 6), "ts": time.time()}
    print(json.dumps(row))        # send to your warehouse / observability pipeline instead

Check each provider's token-accounting semantics (for example, whether cached tokens are included in the input total) against current docs, and adjust the adapters.

Budgets and guardrails

  • Per-request caps: max_tokens / max_output_tokens, input size limits, step caps for agents.
  • Per-user and per-tenant quotas: daily token or cost limits enforced in your backend.
  • Per-feature budgets with alerts at 50%, 80% and 100%.
  • Provider-side limits: spend limits and alerts in each console, separate projects per product.
  • Anomaly alerts: cost per request or per tenant jumping versus the trailing average.

Unit economics

Translate tokens into business terms: cost per support conversation resolved, per lead enriched, per product description published, per active user per month. Compare with the value created (time saved, revenue, conversion). This is how you decide where to use bigger models, where caching or batch are worth engineering time, and how to price AI features.

Worked example: an agency billing clients for AI usage

A UK-and-Pakistan agency runs AI workflows for 25 clients. With the ledger tagged by tenant_id and feature, finance produces a monthly per-client AI cost report, adds a margin, and invoices it. They also discovered that one client's "weekly report" feature cost ten times more than similar clients (illustrative) because it attached a 60-page PDF every run; caching and page selection fixed it.

Forecasting next month

A simple forecast is enough for most teams: take the last four weeks of ledger data per feature, compute cost per outcome and outcomes per week, and project forward with known changes (a new client onboarding, a marketing campaign, a model switch). Add a buffer for growth and for price changes. Share the forecast with finance alongside the budget alerts, and compare forecast to actual each month; the gap tells you which assumptions to fix.

Pitfalls

  • Relying only on the provider invoice.
  • Hard-coded prices that silently go stale.
  • Ignoring reasoning tokens and cache writes.
  • No per-tenant limits, so one customer's loop becomes everyone's bill.

Measuring success

Share of spend attributed to a feature and tenant (target near 100%), forecast accuracy, number of budget alerts acted on, and cost per business outcome over time.

Key takeaways

  • Provider invoices show totals; you need your own attribution by feature, tenant and change.
  • Record usage per call, including cached, cache-write and reasoning tokens, with prices from configuration.
  • Reconcile estimates monthly with provider usage and billing reports.
  • Enforce budgets per request, user, tenant and feature, with anomaly alerts.
  • Express AI cost as unit economics tied to business outcomes.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Your invoice doubled but traffic didn't. What data helps most to find the cause?
  2. Why load prices from configuration with an effective date?
  3. Which tokens are easy to overlook but billed?

Put it into practice

Add a per-call cost ledger with feature and tenant tags to one application, set one budget alert, and produce a cost-per-outcome figure for one feature.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.