---
title: "When NOT to fine-tune: the customization ladder"
description: "The most expensive mistake in custom AI Teams often reach for fine-tuning because it sounds like \"making the model ours\". In 2026 that instinct is…"
url: https://optimizeall.com/learn/fine-tuning-and-custom-models/when-not-to-fine-tune
updated: 2026-10-05
---

Fine-Tuning, Distillation and Custom Models · Decide before you tune · lesson 1 of 16 · 15 min

# When NOT to fine-tune: the customization ladder

## The most expensive mistake in custom AI

Teams often reach for fine-tuning because it sounds like "making the model ours". In 2026 that instinct is frequently wrong. Frontier and strong open-weight models follow instructions and formats far better than earlier generations, and prompt caching, long context and retrieval make prompting cheaper and more powerful. OpenAI itself cited this when it began winding down its self-serve fine-tuning platform in 2026, noting that newer base models follow instructions and formats much better and that prompt-based approaches were often cheaper and faster. Fine-tuning is still valuable, but for **specific** problems. This lesson gives you the ladder to climb first and the tests that tell you fine-tuning is justified.

## What fine-tuning is good (and bad) at

Fine-tuning adjusts a model's weights using examples. It is good at changing **behavior**:

- Consistent **format and style** (a brand voice, a strict output structure) without long prompts.
- **Narrow classification or extraction** where a small, fast, cheap model must match a big model's accuracy.
- **Distillation**: teaching a small model to imitate a large model on one task, cutting cost and latency.
- **Domain-specific reasoning patterns** that prompting cannot fully elicit, especially with preference or reinforcement fine-tuning and a clear grader.
- **Latency and privacy**: a small fine-tuned model can run locally or on-device.

It is poor at:

- **Adding or updating knowledge.** Facts learned in fine-tuning are hard to update, easy to hallucinate around and impossible to cite. Use retrieval (RAG) for knowledge.
- **Fixing unclear requirements.** If humans disagree on the right answer, training will not resolve it.
- **Quick iteration.** Every change means new data, a new training run and a new evaluation.

## The customization ladder

Climb one rung at a time and **measure** on the same evaluation set at each step.

| Rung | What you do | Typical effort |
|---|---|---|
| 1. Better prompt | Clear instructions, role, constraints, output schema, examples of edge cases | Hours |
| 2. Few-shot examples | 3–10 curated examples in the prompt (cacheable) | Hours |
| 3. Structured output + tools | Schemas, function calling, deterministic code for logic | Days |
| 4. Retrieval (RAG) | Ground answers in your documents | Days to weeks |
| 5. Model choice and routing | Try a stronger or cheaper model; route by task | Days |
| 6. Fine-tuning | SFT, LoRA, preference or reinforcement fine-tuning, distillation | Weeks, plus ongoing maintenance |

Most teams find their bar is met at rungs 1–5. Fine-tune when the evidence shows the remaining gap is behavioural and valuable to close.

## Five signals that fine-tuning is justified

1. **Your prompt is long and fragile**, and the same instructions are repeated for every request at high volume.
2. **A small model almost passes** and you need its cost, latency or on-device deployment.
3. **You have (or can create) hundreds to thousands of high-quality examples** that encode the behavior.
4. **You can measure success automatically or with a clear rubric**, before and after.
5. **The behavior is stable**: the task definition will not change every month.

If fewer than three are true, keep climbing the ladder.

## Worked example: a brand-voice assistant for a UAE retailer

A Dubai fashion retailer wants product descriptions in its voice across English and Arabic.

- **Rung 1–2:** a style guide in the system prompt plus six example descriptions. Editors approve 70% without edits (illustrative).
- **Rung 3:** a JSON schema for fields (title, bullets, care, sizing note) removes formatting errors.
- **Rung 5:** a stronger model raises approval to 84% but costs more per description.
- **Signals:** 40,000 descriptions a year, a 2,000-description approved archive, a clear rubric, a stable voice. Four of five signals are true.
- **Decision:** distil the stronger model's style into a small open model with LoRA, evaluate against the rubric, and run it at a fraction of the cost. Keep the stronger model for new categories.

## Hands-on: the "should we fine-tune?" worksheet

```yaml
task: product descriptions in brand voice (EN/AR)
current_best:
  rung: 5
  setup: strong API model + style guide + 6 examples + JSON schema
  eval_set: 150 items, rubric 1-5 on voice, accuracy, format; editor approval rate
  score: approval 84%, mean rubric 4.2
target: approval >= 85% at <= 30% of current cost per item, p95 latency < 3s
gap_type: cost + latency (behavior already acceptable)   # knowledge gap? -> RAG, not fine-tuning
signals:
  long_fragile_prompt: true
  small_model_almost_passes: true     # 71% approval with prompt alone
  examples_available: 2000 approved descriptions
  measurable: true
  stable_behavior: true
decision: fine-tune (LoRA distillation into small open model), re-evaluate in 3 weeks
owner: ai-lead@retailer.example
```

## Pitfalls

- Fine-tuning to add product facts or policies (use RAG; facts change).
- Skipping the baseline, so you cannot prove the fine-tune helped.
- Underestimating maintenance: every base-model upgrade may require retraining.
- Assuming hosted fine-tuning will always be available for your chosen provider (platform offerings change; see Module 5).

## How to measure success

You have a completed worksheet showing each rung's score on the same evaluation set, the specific gap fine-tuning addresses, and a target that makes the effort worthwhile.

## Video lecture: When NOT to fine-tune: the customization ladder

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. When NOT to fine-tune
2. Analogy: a new employee
3. 2026 context
4. Good at: behavior
5. Bad at
6. The customization ladder
7. Five signals
8. Worked example: Dubai retailer (illustrative)
9. Hands-on: the worksheet
10. Simple example: bakery menu
11. The meeting test
12. FAQ
13. Try this now
14. Watch me do it
15. Recap

## Lecture transcript

### When NOT to fine-tune

A company spent three months fine-tuning a model to answer questions about its products. It launched, and within two weeks half the answers were outdated because prices and policies had changed. The fix was a retrieval system that took a week. In this lesson you will learn when not to fine-tune, the customization ladder to climb first, and the five signals that tell you fine-tuning is genuinely worth it.

### Analogy: a new employee

Here is an analogy. Prompting is like giving a talented new employee a clear brief. Retrieval is giving them access to the company handbook and files. Fine-tuning is sending them on a months-long training course that changes their habits. You would not send someone on a training course to learn this week's price list. You would give them the price list. But if you need them to write in the house style every single day without reminders, training starts to make sense.

### 2026 context

Here is the context for twenty twenty-six. Modern models follow instructions and formats far better than earlier generations, and prompt caching, long context and retrieval make prompting cheaper and more capable. When OpenAI began winding down its self-serve fine-tuning platform this year, it pointed to exactly that: newer models follow formats much better, and prompt-based approaches are often cheaper and faster. Fine-tuning is still powerful, but for specific problems, not as a default.

### Good at: behavior

So what is fine-tuning good at? Changing behavior. Consistent format and style, like a brand voice, without long prompts. Narrow classification or extraction where a small cheap model must match a big one. Distillation, teaching a small model to imitate a large one on a single task. Some domain reasoning patterns with a clear grader. And latency and privacy, because a small tuned model can run locally.

### Bad at

And what is it bad at? Adding knowledge. Facts learned during fine-tuning are hard to update, easy to hallucinate around and impossible to cite. That is retrieval's job. It also cannot fix unclear requirements. If your experts disagree on the right answer, training will not settle it. And it is slow to iterate, because every change means new data, a new run and a new evaluation.

### The customization ladder

Now the ladder. Rung one: a better prompt, with clear instructions, constraints and an output schema. Rung two: a handful of curated examples, which caching makes cheap. Rung three: structured output, tools and plain code for deterministic logic. Rung four: retrieval to ground answers in your documents. Rung five: model choice and routing, a stronger model or a cheaper one per task. Rung six: fine-tuning. Climb one rung at a time, and measure each one on the same evaluation set.

### Five signals

How do you know you have reached rung six for real? Look for five signals. One, your prompt is long and fragile and repeated at high volume. Two, a small model almost passes and you need its cost, speed or on-device deployment. Three, you have hundreds to thousands of high-quality examples. Four, you can measure success automatically or with a clear rubric. Five, the behavior is stable. If fewer than three are true, keep climbing the ladder.

### Worked example: Dubai retailer (illustrative)

A worked example, with illustrative numbers. A Dubai fashion retailer wants product descriptions in its brand voice, in English and Arabic. A style guide and six examples get seventy percent editor approval. A JSON schema removes formatting errors. A stronger model reaches eighty-four percent, but costs more. They write forty thousand descriptions a year, have two thousand approved examples, a clear rubric and a stable voice. Four of five signals. So they distil the stronger model's style into a small open model and keep the big one for new categories.

### Hands-on: the worksheet

Your hands-on tool is a one-page worksheet. The task, your current best setup and its score, the target, the type of gap, whether it is knowledge, behavior, cost or latency, the five signals, the decision and an owner. If the gap is knowledge, the answer is retrieval, not fine-tuning. And remember the pitfalls: no baseline means no proof, every base-model upgrade may mean retraining, and hosted fine-tuning offerings change, as we will see in module five.

### Simple example: bakery menu

A simple example. A small bakery in Lahore wants its AI assistant to answer questions about today's menu and opening hours. Someone suggests fine-tuning. But the menu changes daily and the hours change during Ramadan. That is knowledge, and it changes often. The right answer is a short prompt plus a retrieval step that reads the current menu file. It took an afternoon, and when the menu changes, they edit one file. No training run needed.

### The meeting test

Let me give you a quick test you can run in a meeting. When someone proposes fine-tuning, ask three questions. What is the score of our best prompted setup on our evaluation set today? What specific gap remains: knowledge, behavior, cost or latency? And who will maintain the model when the base model changes next year? If nobody can answer the first question, the project is not ready. If the gap is knowledge, the answer is retrieval. And if nobody owns maintenance, the fine-tune will quietly rot.

### FAQ

A question from executives: are we falling behind if we do not fine-tune? No. Many of the strongest AI products in production run on well-designed prompts, retrieval and tools, with careful evaluation, and no fine-tuning at all. Fine-tuning is a tool for specific gaps, not a badge of sophistication. A second question: if we fine-tune, will we have to redo it when the base model changes? Usually yes. Plan for retraining, and keep your data and evaluation ready so it takes days, not months.

### Try this now

Try this now. Pick one AI task your team wants to improve and write down its current score on ten real examples, using your best prompt. Then write one sentence describing the remaining gap and label it: knowledge, behavior, cost or latency. If it says knowledge, stop and plan retrieval instead. If it says behavior, cost or latency, keep it for the worksheet.

### Watch me do it

Watch me do it. A product manager asks me to fine-tune a model for support replies. I open the worksheet. Current best setup: a strong API model with a good prompt. I run it on the fifty-item evaluation set with the rubric: mean score four point one. Target: four point three, at a lower cost per reply. Now the gap: I read the ten lowest-scoring replies. Six are wrong about this month's shipping promotion. That is knowledge, so I add retrieval over the promotions page and rerun: mean four point three. Target met on quality. Remaining gap: cost, since volume is one hundred thousand replies a month. I check the signals: long repeated prompt, yes; a small model almost passes, I test one, it scores three point nine; two thousand approved replies exist; rubric exists; tone is stable. Five of five. My written decision: retrieval now, and a distillation project for cost, starting next sprint.

### Recap

Recap. Fine-tuning changes behavior, not knowledge. Climb the ladder: prompt, examples, structure, retrieval, model choice, and only then fine-tuning. Look for the five signals, and measure every rung on the same evaluation set. Your next step: fill in the worksheet for one real task and decide, with evidence, whether fine-tuning is justified.

## Key takeaways

- Fine-tuning changes behavior (format, style, narrow skills); it is poor at adding knowledge
- Climb the ladder first: prompt, few-shot, structure/tools, RAG, model choice and routing
- Fine-tune when a long prompt, a nearly-passing small model, good examples, measurability and stability align
- Measure every rung on the same evaluation set
- Budget for ongoing maintenance and changing platform offerings

## Try it

Complete the fine-tuning worksheet for one real task: record the score at each rung you have tried and decide whether fine-tuning is justified.

- [Next: The customization methods map: SFT, PEFT, preference, RFT and distillation](https://optimizeall.com/learn/fine-tuning-and-custom-models/customization-methods-map)
- [All lessons of Fine-Tuning, Distillation and Custom Models](https://optimizeall.com/learn/fine-tuning-and-custom-models)
