Fine-Tuning, Distillation and Custom ModelsDecide before you tune · Lesson 2 of 16

The customization methods map: SFT, PEFT, preference, RFT and distillation

Article · 15 min · 9 min lecture

Video lecture

The customization methods map: SFT, PEFT, preference, RFT and distillation

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

The methods map

  • One vocabulary
  • A decision flow
  • Data formats per method

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

One vocabulary for the whole course

"Fine-tuning" covers several distinct techniques. Mixing them up leads to the wrong data, the wrong tools and the wrong expectations. Here is the map you will use for the rest of the course.

The methods

MethodWhat you provideWhat it changesTypical use
Continued pre-trainingLarge amounts of raw domain textGeneral familiarity with a domain's languageRare for businesses; expensive; e.g. a new language or very specialized corpus
Supervised fine-tuning (SFT)Input → ideal output pairs (often chat transcripts)Imitates the demonstrated behaviorFormat, style, classification, extraction, tool-use patterns
Parameter-efficient fine-tuning (PEFT), e.g. LoRA/QLoRASame data as SFT (or preference data)Trains small adapter weights instead of all weightsAlmost all practical open-model fine-tuning
Preference optimization (RLHF, DPO and variants)Pairs: a prompt with a preferred and a rejected responseShifts the model toward preferred responsesTone, helpfulness, safety, subjective quality
Reinforcement fine-tuning (RFT, e.g. GRPO)Prompts plus a grader (code or model) that scores outputsRewards outputs that score higherTasks with checkable answers: maths, code, extraction, policy compliance
DistillationOutputs from a strong "teacher" model on your taskTrains a smaller "student" to imitate the teacherCutting cost and latency while keeping quality on one task
Embedding fine-tuningQuery → relevant passage pairs (plus hard negatives)Improves retrieval similarity for your domainBetter RAG retrieval

Notes:

  • PEFT is a how (which weights you train), not a what (which objective). You can do SFT with LoRA, DPO with LoRA, and GRPO with LoRA.
  • Distillation is usually implemented as SFT on teacher outputs (sometimes with logit-level distillation when you control both models).
  • Model merging (combining the weights of several fine-tunes) and multi-LoRA serving are deployment techniques you will meet in Module 6.

Hosted vs self-managed

PathProsCons
Hosted fine-tuning APIs (cloud providers, model vendors, inference platforms)No GPU operations; integrated servingLimited to supported models and methods; data goes to the provider; offerings change (for example, OpenAI began winding down its self-serve platform in 2026)
Self-managed on open weights (TRL, PEFT, Unsloth, Axolotl on your GPUs or rented cloud GPUs)Full control, any open model, data stays where you choose, portable adaptersYou own the pipeline, compute and evaluation

A common 2026 pattern: prototype with the ladder on API models, then distil or fine-tune an open-weight model you control for the high-volume path.

Choosing a method: a decision flow

  1. Is the gap knowledge? Use RAG. Stop.
  2. Can you write ideal outputs? Yes → SFT (with LoRA).
  3. Is quality subjective, and easier to compare than to write? Collect preference pairs → DPO after SFT.
  4. Is correctness checkable by code or a reliable grader? → RFT/GRPO, usually after SFT.
  5. Is the goal cheaper/faster at similar quality? → Distil a strong model's outputs into a small model (SFT on teacher outputs).
  6. Is retrieval missing the right passages? → Fine-tune the embedding model (or add a reranker first).

Worked example: one company, four methods

A Pakistani fintech app runs customer support in English, Urdu and Roman Urdu.

  • Intent routing (18 intents): SFT with LoRA on a 3B model using 6,000 labeled messages. Fast and cheap.
  • Reply tone ("respectful, concise, no promises about approvals"): agents are better at choosing between two drafts than writing perfect ones, so they collect preference pairs and run DPO on top of the SFT model.
  • Fee calculations in replies: a grader checks that quoted fees match the fee schedule; RFT improves accuracy on that check.
  • Help-center search: Roman Urdu queries missed English articles, so they fine-tune the embedding model on query-article pairs.

Each method matched a specific, measurable gap.

Hands-on: the data you need for each method

{"messages": [{"role": "system", "content": "Classify the intent."}, {"role": "user", "content": "mera card block ho gaya"}, {"role": "assistant", "content": "card_blocked"}]}
{"prompt": [{"role": "user", "content": "Why was my loan rejected?"}], "chosen": [{"role": "assistant", "content": "I'm sorry to hear that. Decisions consider several factors; I can share how to request the reasons and next steps."}], "rejected": [{"role": "assistant", "content": "Your credit score is bad. Try again later."}]}
{"prompt": [{"role": "user", "content": "What is the fee for a PKR 50,000 transfer?"}], "expected_fee": "PKR 0"}
{"anchor": "bill payment nahi ho rahi", "positive": "Troubleshooting failed bill payments: check biller ID and account balance..."}

These are, in order: SFT (conversational), DPO (preference), GRPO/RFT (prompt plus a field your grader uses) and embedding fine-tuning (anchor/positive). The fee value above is illustrative.

Pitfalls

  • Choosing a method by hype ("we need RLHF") rather than by gap type.
  • Running DPO before the model can do the task at all (do SFT first).
  • Building RFT without a reliable grader; the model will learn to exploit a weak one.
  • Assuming hosted methods are available for every model.

How to measure success

For each gap, you can name the method, the data format, the evaluation metric and the reason the method fits the gap.

Key takeaways

  • SFT imitates demonstrations; PEFT/LoRA is how you train cheaply, not a separate objective
  • Preference optimization (RLHF/DPO) fits subjective quality; RFT/GRPO fits checkable correctness
  • Distillation trains a small student on a strong teacher's outputs to cut cost and latency
  • Embedding fine-tuning improves retrieval, not generation
  • Pick the method by gap type; hosted vs self-managed is a separate decision

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Agents find it easier to choose the better of two replies than to write a perfect reply. Which method fits?
  2. Which statement about LoRA is correct?
  3. Your correctness can be checked by code (e.g. computed fees must match a schedule). Which method is designed for this?

Put it into practice

For one product, list three performance gaps and map each to a method, data format and evaluation metric using the decision flow.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.