Fine-Tuning, Distillation and Custom ModelsDecide before you tune · Lesson 1 of 16

When NOT to fine-tune: the customization ladder

Article · 15 min · 9 min lecture

Video lecture

When NOT to fine-tune: the customization ladder

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

When NOT to fine-tune

  • The customization ladder
  • Five signals that justify it

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The most expensive mistake in custom AI

Teams often reach for fine-tuning because it sounds like "making the model ours". In 2026 that instinct is frequently wrong. Frontier and strong open-weight models follow instructions and formats far better than earlier generations, and prompt caching, long context and retrieval make prompting cheaper and more powerful. OpenAI itself cited this when it began winding down its self-serve fine-tuning platform in 2026, noting that newer base models follow instructions and formats much better and that prompt-based approaches were often cheaper and faster. Fine-tuning is still valuable, but for specific problems. This lesson gives you the ladder to climb first and the tests that tell you fine-tuning is justified.

What fine-tuning is good (and bad) at

Fine-tuning adjusts a model's weights using examples. It is good at changing behavior:

  • Consistent format and style (a brand voice, a strict output structure) without long prompts.
  • Narrow classification or extraction where a small, fast, cheap model must match a big model's accuracy.
  • Distillation: teaching a small model to imitate a large model on one task, cutting cost and latency.
  • Domain-specific reasoning patterns that prompting cannot fully elicit, especially with preference or reinforcement fine-tuning and a clear grader.
  • Latency and privacy: a small fine-tuned model can run locally or on-device.

It is poor at:

  • Adding or updating knowledge. Facts learned in fine-tuning are hard to update, easy to hallucinate around and impossible to cite. Use retrieval (RAG) for knowledge.
  • Fixing unclear requirements. If humans disagree on the right answer, training will not resolve it.
  • Quick iteration. Every change means new data, a new training run and a new evaluation.

The customization ladder

Climb one rung at a time and measure on the same evaluation set at each step.

RungWhat you doTypical effort
1. Better promptClear instructions, role, constraints, output schema, examples of edge casesHours
2. Few-shot examples3–10 curated examples in the prompt (cacheable)Hours
3. Structured output + toolsSchemas, function calling, deterministic code for logicDays
4. Retrieval (RAG)Ground answers in your documentsDays to weeks
5. Model choice and routingTry a stronger or cheaper model; route by taskDays
6. Fine-tuningSFT, LoRA, preference or reinforcement fine-tuning, distillationWeeks, plus ongoing maintenance

Most teams find their bar is met at rungs 1–5. Fine-tune when the evidence shows the remaining gap is behavioural and valuable to close.

Five signals that fine-tuning is justified

  1. Your prompt is long and fragile, and the same instructions are repeated for every request at high volume.
  2. A small model almost passes and you need its cost, latency or on-device deployment.
  3. You have (or can create) hundreds to thousands of high-quality examples that encode the behavior.
  4. You can measure success automatically or with a clear rubric, before and after.
  5. The behavior is stable: the task definition will not change every month.

If fewer than three are true, keep climbing the ladder.

Worked example: a brand-voice assistant for a UAE retailer

A Dubai fashion retailer wants product descriptions in its voice across English and Arabic.

  • Rung 1–2: a style guide in the system prompt plus six example descriptions. Editors approve 70% without edits (illustrative).
  • Rung 3: a JSON schema for fields (title, bullets, care, sizing note) removes formatting errors.
  • Rung 5: a stronger model raises approval to 84% but costs more per description.
  • Signals: 40,000 descriptions a year, a 2,000-description approved archive, a clear rubric, a stable voice. Four of five signals are true.
  • Decision: distil the stronger model's style into a small open model with LoRA, evaluate against the rubric, and run it at a fraction of the cost. Keep the stronger model for new categories.

Hands-on: the "should we fine-tune?" worksheet

task: product descriptions in brand voice (EN/AR)
current_best:
  rung: 5
  setup: strong API model + style guide + 6 examples + JSON schema
  eval_set: 150 items, rubric 1-5 on voice, accuracy, format; editor approval rate
  score: approval 84%, mean rubric 4.2
target: approval >= 85% at <= 30% of current cost per item, p95 latency < 3s
gap_type: cost + latency (behavior already acceptable)   # knowledge gap? -> RAG, not fine-tuning
signals:
  long_fragile_prompt: true
  small_model_almost_passes: true     # 71% approval with prompt alone
  examples_available: 2000 approved descriptions
  measurable: true
  stable_behavior: true
decision: fine-tune (LoRA distillation into small open model), re-evaluate in 3 weeks
owner: ai-lead@retailer.example

Pitfalls

  • Fine-tuning to add product facts or policies (use RAG; facts change).
  • Skipping the baseline, so you cannot prove the fine-tune helped.
  • Underestimating maintenance: every base-model upgrade may require retraining.
  • Assuming hosted fine-tuning will always be available for your chosen provider (platform offerings change; see Module 5).

How to measure success

You have a completed worksheet showing each rung's score on the same evaluation set, the specific gap fine-tuning addresses, and a target that makes the effort worthwhile.

Key takeaways

  • Fine-tuning changes behavior (format, style, narrow skills); it is poor at adding knowledge
  • Climb the ladder first: prompt, few-shot, structure/tools, RAG, model choice and routing
  • Fine-tune when a long prompt, a nearly-passing small model, good examples, measurability and stability align
  • Measure every rung on the same evaluation set
  • Budget for ongoing maintenance and changing platform offerings

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A support team wants the model to know this month's updated refund policy. What is the best approach?
  2. Which situation most strongly justifies fine-tuning?
  3. Why run every rung of the ladder on the same evaluation set?

Put it into practice

Complete the fine-tuning worksheet for one real task: record the score at each rung you have tried and decide whether fine-tuning is justified.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.