---
title: "The customization methods map: SFT, PEFT, preference, RFT…"
description: "One vocabulary for the whole course \"Fine-tuning\" covers several distinct techniques. Mixing them up leads to the wrong data, the wrong tools and the…"
url: https://optimizeall.com/learn/fine-tuning-and-custom-models/customization-methods-map
updated: 2026-10-05
---

Fine-Tuning, Distillation and Custom Models · Decide before you tune · lesson 2 of 16 · 15 min

# The customization methods map: SFT, PEFT, preference, RFT and distillation

## One vocabulary for the whole course

"Fine-tuning" covers several distinct techniques. Mixing them up leads to the wrong data, the wrong tools and the wrong expectations. Here is the map you will use for the rest of the course.

## The methods

| Method | What you provide | What it changes | Typical use |
|---|---|---|---|
| **Continued pre-training** | Large amounts of raw domain text | General familiarity with a domain's language | Rare for businesses; expensive; e.g. a new language or very specialized corpus |
| **Supervised fine-tuning (SFT)** | Input → ideal output pairs (often chat transcripts) | Imitates the demonstrated behavior | Format, style, classification, extraction, tool-use patterns |
| **Parameter-efficient fine-tuning (PEFT), e.g. LoRA/QLoRA** | Same data as SFT (or preference data) | Trains small adapter weights instead of all weights | Almost all practical open-model fine-tuning |
| **Preference optimization (RLHF, DPO and variants)** | Pairs: a prompt with a preferred and a rejected response | Shifts the model toward preferred responses | Tone, helpfulness, safety, subjective quality |
| **Reinforcement fine-tuning (RFT, e.g. GRPO)** | Prompts plus a **grader** (code or model) that scores outputs | Rewards outputs that score higher | Tasks with checkable answers: maths, code, extraction, policy compliance |
| **Distillation** | Outputs from a strong "teacher" model on your task | Trains a smaller "student" to imitate the teacher | Cutting cost and latency while keeping quality on one task |
| **Embedding fine-tuning** | Query → relevant passage pairs (plus hard negatives) | Improves retrieval similarity for your domain | Better RAG retrieval |

Notes:

- PEFT is a **how** (which weights you train), not a **what** (which objective). You can do SFT with LoRA, DPO with LoRA, and GRPO with LoRA.
- Distillation is usually implemented as SFT on teacher outputs (sometimes with logit-level distillation when you control both models).
- **Model merging** (combining the weights of several fine-tunes) and **multi-LoRA serving** are deployment techniques you will meet in Module 6.

## Hosted vs self-managed

| Path | Pros | Cons |
|---|---|---|
| **Hosted fine-tuning APIs** (cloud providers, model vendors, inference platforms) | No GPU operations; integrated serving | Limited to supported models and methods; data goes to the provider; offerings change (for example, OpenAI began winding down its self-serve platform in 2026) |
| **Self-managed on open weights** (TRL, PEFT, Unsloth, Axolotl on your GPUs or rented cloud GPUs) | Full control, any open model, data stays where you choose, portable adapters | You own the pipeline, compute and evaluation |

A common 2026 pattern: prototype with the ladder on API models, then distil or fine-tune an open-weight model you control for the high-volume path.

## Choosing a method: a decision flow

1. **Is the gap knowledge?** Use RAG. Stop.
2. **Can you write ideal outputs?** Yes → SFT (with LoRA).
3. **Is quality subjective, and easier to compare than to write?** Collect preference pairs → DPO after SFT.
4. **Is correctness checkable by code or a reliable grader?** → RFT/GRPO, usually after SFT.
5. **Is the goal cheaper/faster at similar quality?** → Distil a strong model's outputs into a small model (SFT on teacher outputs).
6. **Is retrieval missing the right passages?** → Fine-tune the embedding model (or add a reranker first).

## Worked example: one company, four methods

A Pakistani fintech app runs customer support in English, Urdu and Roman Urdu.

- **Intent routing** (18 intents): SFT with LoRA on a 3B model using 6,000 labeled messages. Fast and cheap.
- **Reply tone** ("respectful, concise, no promises about approvals"): agents are better at choosing between two drafts than writing perfect ones, so they collect preference pairs and run DPO on top of the SFT model.
- **Fee calculations in replies:** a grader checks that quoted fees match the fee schedule; RFT improves accuracy on that check.
- **Help-center search:** Roman Urdu queries missed English articles, so they fine-tune the embedding model on query-article pairs.

Each method matched a specific, measurable gap.

## Hands-on: the data you need for each method

```jsonl
{"messages": [{"role": "system", "content": "Classify the intent."}, {"role": "user", "content": "mera card block ho gaya"}, {"role": "assistant", "content": "card_blocked"}]}
```

```jsonl
{"prompt": [{"role": "user", "content": "Why was my loan rejected?"}], "chosen": [{"role": "assistant", "content": "I'm sorry to hear that. Decisions consider several factors; I can share how to request the reasons and next steps."}], "rejected": [{"role": "assistant", "content": "Your credit score is bad. Try again later."}]}
```

```jsonl
{"prompt": [{"role": "user", "content": "What is the fee for a PKR 50,000 transfer?"}], "expected_fee": "PKR 0"}
```

```jsonl
{"anchor": "bill payment nahi ho rahi", "positive": "Troubleshooting failed bill payments: check biller ID and account balance..."}
```

These are, in order: SFT (conversational), DPO (preference), GRPO/RFT (prompt plus a field your grader uses) and embedding fine-tuning (anchor/positive). The fee value above is illustrative.

## Pitfalls

- Choosing a method by hype ("we need RLHF") rather than by gap type.
- Running DPO before the model can do the task at all (do SFT first).
- Building RFT without a reliable grader; the model will learn to exploit a weak one.
- Assuming hosted methods are available for every model.

## How to measure success

For each gap, you can name the method, the data format, the evaluation metric and the reason the method fits the gap.

## Video lecture: The customization methods map: SFT, PEFT, preference, RFT and distillation

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. The methods map
2. Five teaching styles
3. Map, part 1
4. Map, part 2
5. Hosted vs self-managed
6. Decision flow
7. One company, four methods
8. Data shapes
9. Pitfalls
10. Simple example: recipe app
11. Sequencing
12. FAQ: combine methods?
13. Try this now
14. Watch me do it
15. Recap

## Lecture transcript

### The methods map

Ask five engineers what fine-tuning means and you will get five answers. One means training on examples, one means LoRA, one means RLHF, one means distillation, and one means making the search better. They are all different techniques with different data, tools and results. In this lesson you will get one clear map of customization methods, a decision flow to pick the right one, and the exact data format each needs.

### Five teaching styles

Here is a way to remember the methods. SFT is teaching by showing: here is the task, here is how I would do it. Preference optimization is teaching by comparing: this answer is better than that one. Reinforcement fine-tuning is teaching by scoring: try many times, and I will mark each attempt. Distillation is apprenticeship: watch the master and copy. And embedding fine-tuning is training a librarian to find the right book. Five teaching styles for five kinds of learning.

### Map, part 1

Let's walk the map. Continued pre-training feeds large amounts of raw domain text, and it is rare and expensive for businesses. Supervised fine-tuning, SFT, gives input and ideal output pairs, and the model imitates the demonstrated behavior. Parameter-efficient fine-tuning, like LoRA and QLoRA, trains small adapter weights instead of the whole model. Important: that is a how, not a what. You can do SFT with LoRA, preference training with LoRA, and reinforcement training with LoRA.

### Map, part 2

Continuing. Preference optimization, which includes RLHF and DPO, uses pairs: a prompt, a preferred response and a rejected one, and nudges the model toward preferred answers. Great for tone and subjective quality. Reinforcement fine-tuning, with methods like GRPO, uses prompts plus a grader that scores outputs, perfect when correctness can be checked. Distillation trains a small student to imitate a strong teacher on your task. And embedding fine-tuning improves retrieval similarity, not generation.

### Hosted vs self-managed

Separate from the method is where you run it. Hosted fine-tuning APIs from cloud providers, model vendors and inference platforms mean no GPU operations and integrated serving, but you are limited to their models and methods, your data goes to them, and offerings change. OpenAI, for example, began winding down its self-serve fine-tuning platform in twenty twenty-six. Self-managed training on open weights, with tools like TRL, PEFT, Unsloth and Axolotl, gives full control and portable adapters, but you own the pipeline.

### Decision flow

Here is the decision flow. Is the gap knowledge? Use retrieval and stop. Can you write ideal outputs? Use SFT with LoRA. Is quality subjective and easier to compare than to write? Collect preference pairs and run DPO after SFT. Is correctness checkable by code or a reliable grader? Use reinforcement fine-tuning, usually after SFT. Is the goal cheaper and faster at similar quality? Distil. Is retrieval missing the right passages? Fine-tune the embedding model, or try a reranker first.

### One company, four methods

One company, four methods. A Pakistani fintech app supports customers in English, Urdu and Roman Urdu. For routing eighteen intents, they use SFT with LoRA on a small model and six thousand labeled messages. For reply tone, agents find it easier to pick the better of two drafts, so they run DPO on top. For fee amounts in replies, a grader checks quotes against the fee schedule, so they use reinforcement fine-tuning. And because Roman Urdu searches missed English help articles, they fine-tune the embedding model.

### Data shapes

Each method needs a different data shape, and the lesson shows all four. SFT uses chat messages ending in the ideal assistant reply. DPO uses a prompt, a chosen response and a rejected one. Reinforcement fine-tuning uses a prompt plus whatever fields your grader needs, like the expected fee. And embedding fine-tuning uses an anchor, the query, and a positive, the passage that should match. Getting the shape right on day one saves weeks.

### Pitfalls

Four pitfalls. Choosing a method by hype, like we need RLHF, instead of by gap type. Running DPO before the model can do the task at all, so do SFT first. Building reinforcement fine-tuning without a reliable grader, because the model will learn to exploit a weak one. And assuming a hosted provider supports every method for every model. Check the docs.

### Simple example: recipe app

A simple example. A recipe app wants three improvements. First, always output recipes in the same structured format: that is SFT with examples of the format. Second, users prefer friendlier, shorter tips, and it is easier to pick the better tip than to write the perfect one: that is preference data and DPO. Third, when a user searches for quick vegan dinner, results often miss the right recipes: that is an embedding problem. Three gaps, three different methods, and none of them is simply more fine-tuning.

### Sequencing

A practical tip on sequencing. Most successful projects follow the same order: first SFT to teach the task and format, then, only if needed, preference optimization for subjective quality, or reinforcement fine-tuning for checkable correctness. Distillation is usually SFT with teacher outputs as the data. And embedding fine-tuning runs on its own track, because it improves retrieval rather than generation. Skipping SFT and jumping straight to DPO or GRPO usually gives weak results, because the model cannot yet do the basic task.

### FAQ: combine methods?

A common question: can we combine methods in one model? Yes, and it is normal. A typical sequence is SFT to teach the task, then DPO for tone, or GRPO for checkable correctness, all with LoRA adapters on the same base model. Separately, you might fine-tune an embedding model for retrieval. Each step needs its own evaluation, and each should earn its place: if DPO does not improve your blind win rate, drop it. More methods are not automatically better.

### Try this now

Try this now. For one product you know, write three sentences that each start with the model currently fails when. Then next to each, write which method from the map fits, and what the training data would look like in one line. If you cannot describe the data, that is the real blocker to solve first.

### Watch me do it

Watch me do it. I take a real product, a telecom self-service app, and list its three biggest AI complaints from support tickets. One: the assistant answers in long paragraphs when customers want short steps. Two: agents prefer a warmer tone than the model uses. Three: searches for eSIM activation do not find the right help article. For the first, I can write ideal short answers, so I mark SFT with LoRA and sketch one line of messages data. For the second, it is easier to compare than to write, so I mark DPO after SFT and sketch a chosen and rejected pair. For the third, retrieval is failing, so I check recall at fifty first: the right article is not even in the candidates, so a reranker will not help. I mark embedding fine-tuning with anchor and positive pairs. Three gaps, three methods, three data formats, written down in ten minutes.

### Recap

Recap. SFT imitates demonstrations. LoRA is how you train cheaply, with any objective. DPO learns from preferences, reinforcement fine-tuning learns from a grader, distillation shrinks a teacher into a student, and embedding fine-tuning improves retrieval. Choose by gap type, then choose hosted or self-managed. Your next step: list three gaps in one product and map each to a method, data format and metric.

## Key takeaways

- SFT imitates demonstrations; PEFT/LoRA is how you train cheaply, not a separate objective
- Preference optimization (RLHF/DPO) fits subjective quality; RFT/GRPO fits checkable correctness
- Distillation trains a small student on a strong teacher's outputs to cut cost and latency
- Embedding fine-tuning improves retrieval, not generation
- Pick the method by gap type; hosted vs self-managed is a separate decision

## Try it

For one product, list three performance gaps and map each to a method, data format and evaluation metric using the decision flow.

- [Previous: When NOT to fine-tune: the customization ladder](https://optimizeall.com/learn/fine-tuning-and-custom-models/when-not-to-fine-tune)
- [Next: Dataset design and curation: quality beats quantity](https://optimizeall.com/learn/fine-tuning-and-custom-models/dataset-design-and-curation)
- [All lessons of Fine-Tuning, Distillation and Custom Models](https://optimizeall.com/learn/fine-tuning-and-custom-models)
