Fine-Tuning, Distillation and Custom Models · Evaluation, hosted platforms, embeddings and cost · lesson 14 of 16 · 15 min
Cost modeling: training, inference and break-even
The question finance will ask
"What does this cost, and when does it pay back?" A fine-tune has one-off costs (data, training, evaluation), recurring costs (hosting, monitoring, retraining) and savings or gains (cheaper inference, better outcomes). This lesson gives you a model you can defend. All numbers below are illustrative placeholders; replace them with your quotes and current price sheets.
Cost components
One-off:
- Data: labeling hours × rate; synthetic generation (teacher tokens or GPU hours); review hours.
- Training compute: hosted price × training tokens, or GPU-hours × hourly rate.
- Evaluation: reviewer hours; judge-model tokens.
- Engineering time to build the pipeline.
Recurring:
- Inference: per-token price (hosted custom models may be priced differently from base models) or GPU hosting.
- Monitoring, drift checks and periodic evaluation.
- Retraining when data drifts or the base model is upgraded or deprecated.
Training compute: two ways to estimate
Hosted (token-based):
training_tokens = examples × avg_tokens_per_example × epochs
training_cost = training_tokens × price_per_training_token
Example (illustrative): 6,000 examples × 350 tokens × 2 epochs = 4.2M training tokens. Multiply by the provider's current training price.
Self-managed (GPU-hours):
gpu_hours = training_tokens ÷ (measured tokens_per_second × 3600)
training_cost = gpu_hours × hourly_rate × (1 + overhead for failed/experimental runs)
Measure tokens per second with a short trial run on the target GPU; do not guess. Budget for several experimental runs (for example, multiply by 3–5 for the whole project).
Inference economics and break-even
The usual motivation is replacing an expensive prompted model with a cheaper fine-tuned one:
saving_per_request = cost_per_request(current) − cost_per_request(fine-tuned)
monthly_saving = saving_per_request × monthly_requests
break_even_months = one_off_costs ÷ (monthly_saving − extra_recurring_costs)
Cost per request for token-priced models:
cost_per_request = input_tokens × input_price + output_tokens × output_price
Remember: a fine-tuned model often needs shorter prompts (no long instructions or few-shot examples), which cuts input tokens too.
Worked example (illustrative numbers)
A marketplace classifies 3 million messages a month.
- Current: strong API model; prompt 900 input tokens (instructions + examples) + 60 message tokens; 5 output tokens. Suppose that comes to $0.0010 per request → $3,000/month.
- Fine-tuned small open model on own GPU: 80 input tokens, 5 output tokens; one GPU server at $900/month all-in handles the volume with headroom.
- One-off: labeling and review $4,000; training experiments $600; evaluation $1,200; engineering $6,000 → $11,800.
- Recurring extras: monitoring and quarterly retraining ≈ $400/month.
- Monthly saving: $3,000 − $900 − $400 = $1,700 → break-even ≈ 7 months.
Then stress-test: what if volume halves? (Monthly saving falls to about $200, so break-even stretches to roughly five years: probably not worth it.) What if the API price halves instead? (The same: saving shrinks to about $200/month.) What if volume doubles? (Saving rises to about $4,700/month if one server still copes, and break-even drops to under 3 months.) What if accuracy gains reduce human escalations? (Add that value explicitly.)
Hands-on: a break-even calculator
# breakeven.py: all inputs are your own estimates; values below are illustrative placeholders
def cost_per_request(inp, out, price_in_per_m, price_out_per_m):
return inp * price_in_per_m / 1e6 + out * price_out_per_m / 1e6
def breakeven(one_off, monthly_requests, current_cpr, new_cpr, new_fixed_monthly, extra_recurring):
monthly_saving = monthly_requests * (current_cpr - new_cpr) - new_fixed_monthly - extra_recurring
return (one_off / monthly_saving) if monthly_saving > 0 else float("inf"), monthly_saving
current = cost_per_request(960, 5, price_in_per_m=1.00, price_out_per_m=4.00) # replace with current prices
new = 0.0 # own GPU: marginal cost ~0, fixed below
for volume in (1_500_000, 3_000_000, 6_000_000):
months, saving = breakeven(one_off=11_800, monthly_requests=volume, current_cpr=current, new_cpr=new,
new_fixed_monthly=900, extra_recurring=400)
print(f"{volume:>9,} req/month: saving ${saving:,.0f}/month, break-even {months:.1f} months")
Beyond cost: value
Fine-tuning can also create value that is not a cost saving: higher accuracy (fewer human escalations), lower latency (better conversion in live chat), on-device operation (new product possibilities), and data residency (market access). Quantify these where you can, for example "each correctly auto-routed ticket saves 3 minutes of agent time".
Pitfalls
- Using list prices from memory instead of current price sheets.
- Ignoring engineering and retraining costs.
- Forgetting prompt caching and batch discounts that make the current API cheaper than it looks.
- Not stress-testing volume and price assumptions.
How to measure success
A one-page cost model with sources for every price, sensitivity analysis for volume and price, and a break-even that still looks acceptable under pessimistic assumptions.
Video lecture: Cost modeling: training, inference and break-even
Lecture coming soon · 15 chapters · about 8 minutes. Read the full transcript below.
- Cost modeling
- Analogy: van vs courier
- Components
- Training compute
- Inference economics
- Break-even (illustrative)
- Stress tests
- Value beyond cost
- Pitfalls
- Simple example (illustrative)
- One-page presentation
- FAQ: what about people costs?
- Try this now
- Watch me do it
- Recap
Lecture transcript
Cost modeling
Sooner or later, someone from finance asks two questions: what does this cost, and when does it pay back? If your answer is it depends, the project stalls. In this lesson you will build a cost model for custom models that you can defend: one-off and recurring costs, two ways to estimate training, inference economics, break-even, and stress tests. Every number I use is illustrative; you will plug in your own.
Analogy: van vs courier
Here is an analogy. Deciding whether to fine-tune for cost is like deciding whether to buy a delivery van instead of paying a courier per parcel. The van has an upfront price, insurance and maintenance. The courier charges per parcel. With a few parcels a week, the courier wins. With hundreds a day, the van pays for itself. Break-even is simply the week when the van's total cost drops below the courier bills you would have paid.
Components
Start with the components. One-off costs: data labeling and review, synthetic generation, training compute, evaluation hours and judge tokens, and engineering time. Recurring costs: inference, whether per token or GPU hosting; monitoring and drift checks; and retraining when data drifts or the base model changes. Hosted custom models may be priced differently from base models, so check.
Training compute
Training compute, two ways. Hosted: training tokens equal examples times average tokens per example times epochs, multiplied by the provider's training price. Six thousand examples of three hundred and fifty tokens for two epochs is four point two million tokens. Self-managed: GPU hours equal training tokens divided by measured tokens per second times thirty-six hundred, times the hourly rate, plus overhead for experiments. Measure throughput with a short trial run, and budget for several runs, not one.
Inference economics
Now inference economics. The usual motivation is replacing an expensive prompted model with a cheaper fine-tuned one. Cost per request is input tokens times input price plus output tokens times output price. Here is a hidden bonus: a fine-tuned model often needs much shorter prompts, because the instructions and examples are baked into the weights. So you save on input tokens as well as the per-token price.
Break-even (illustrative)
Break-even is simple. Monthly saving equals saving per request times monthly requests, minus any new fixed and recurring costs. Break-even months equal one-off costs divided by that monthly saving. In the lesson's illustrative example, a marketplace classifying three million messages a month moves from about three thousand dollars a month on an API to a nine hundred dollar GPU server plus four hundred for monitoring and retraining. With about eleven thousand eight hundred in one-off costs, break-even is roughly seven months.
Stress tests
Now stress-test, because assumptions are where business cases die. If volume halves, the monthly saving falls to about two hundred dollars and break-even stretches to roughly five years, probably not worth it. If the API price halves instead, you get the same result. If volume doubles and one server still copes, break-even drops to under three months. Also remember discounts: prompt caching and batch pricing can make the current setup cheaper than list prices suggest. The calculator in the lesson loops over volumes so you can see the whole picture in seconds.
Value beyond cost
Cost is not the only value. A fine-tune can raise accuracy and reduce human escalations, cut latency and lift conversion in live chat, run on devices to enable new products, or keep data in-country to open a market. Put numbers on these where you can, for example, each correctly auto-routed ticket saves three minutes of agent time, times the number of tickets.
Pitfalls
Four pitfalls. Using prices from memory instead of current price sheets. Ignoring engineering and retraining costs. Forgetting discounts on the current API. And not stress-testing volume and price. A good model is one page, cites a source for every price, shows sensitivity to volume and price, and still makes sense under pessimistic assumptions.
Simple example (illustrative)
A simple example with round, illustrative numbers. A startup spends one thousand dollars a month on API calls for a classification task. A fine-tuned small model would cost two hundred a month to run, plus one hundred for monitoring. One-off costs, data, training and evaluation, come to three thousand. Monthly saving: seven hundred. Break-even: a little over four months. If volume halves, the saving falls to two hundred, and break-even jumps to fifteen months. That second number is the one to show your finance team.
One-page presentation
Present the model to decision-makers on one page. At the top, the recommendation and the break-even under expected assumptions. Below, a small table of three scenarios: pessimistic, expected and optimistic, each with volume, prices and break-even months. Then the value beyond cost, like fewer escalations or faster replies, with your estimate and how you measured it. And finally the sources for every price and the date you checked them. Decision-makers trust numbers they can trace.
FAQ: what about people costs?
A question finance teams ask: what about the cost of people? Include it. Engineering time to build the pipeline, reviewer time for labeling and evaluation, and ongoing time for monitoring and retraining are often larger than the compute bill, especially in the first year. A model that looks cheap on GPU hours can look very different once people's time is included. Honest models include both, and the recommendation still holds.
Try this now
Try this now. Open your provider's current price page and your usage dashboard, and calculate your real cost per request for one AI task, including input and output tokens. Write the date and the source next to the number. Every cost model starts with one honest, sourced number.
Watch me do it
Watch me do it. I open our usage dashboard for the classification feature: one point eight million requests last month, average nine hundred forty input tokens and six output tokens. I open the provider's current price page and note the date. Cost per request: input tokens times the input price plus output tokens times the output price, and I multiply by volume for the monthly bill. For the fine-tuned option, I run a short training trial on a rented GPU and measure tokens per second, which gives GPU hours for the full run; I multiply by four for experiments. I add labeling hours, evaluation hours and engineering days at our loaded rates, and the monthly server and monitoring costs. I put everything into the calculator and run three volume scenarios. Expected case: break-even in about eight months. Pessimistic, half the volume: never within two years. I write both on the summary slide, with sources and dates.
Recap
Recap. Separate one-off and recurring costs. Estimate training by tokens or measured GPU hours. Remember shorter prompts after fine-tuning. Compute break-even, stress-test it, and count value beyond cost. Your next step: fill in the calculator with your real volumes and current prices, and run pessimistic, expected and optimistic scenarios.
Key takeaways
- Separate one-off (data, training, eval, engineering) and recurring (inference, monitoring, retraining) costs
- Estimate training as tokens × epochs × price, or measured GPU-hours × rate with experiment overhead
- Fine-tuned models often need much shorter prompts, cutting input tokens
- Break-even = one-off costs ÷ net monthly saving; stress-test volume and price
- Count value beyond cost: accuracy, latency, residency, new capabilities
Try it
Fill in the break-even calculator with your real volumes and current price sheets, then run pessimistic, expected and optimistic scenarios.