---
title: "Supervised fine-tuning fundamentals | Optimize All Academy"
description: "What actually happens during SFT A language model is trained to predict the next token. Supervised fine-tuning (SFT) continues that training on your…"
url: https://optimizeall.com/learn/fine-tuning-and-custom-models/sft-fundamentals
updated: 2026-10-05
---

Fine-Tuning, Distillation and Custom Models · Supervised fine-tuning and PEFT · lesson 6 of 16 · 16 min

# Supervised fine-tuning fundamentals

## What actually happens during SFT

A language model is trained to predict the next token. **Supervised fine-tuning (SFT)** continues that training on your examples: for each example, the model sees the prompt and the ideal response, and its weights are nudged so the ideal response becomes more likely. Repeat over thousands of examples for one to a few passes (**epochs**), and the model's default behavior shifts toward your demonstrations.

Three details make or break SFT:

1. **Loss masking.** You usually want the model to learn to produce the **response**, not to reproduce your prompts. Computing loss only on completion (assistant) tokens focuses learning where it matters.
2. **The chat template.** Instruction-tuned models expect conversations formatted with specific special tokens. Training and serving must use the **same template**; tools like TRL apply the tokenizer's template for you when data is in conversational format.
3. **Train/serve consistency.** The system prompt, formatting and output style in training must match production exactly.

## Data formats TRL understands

```jsonl
{"messages": [{"role": "system", "content": "Classify the intent."}, {"role": "user", "content": "card blocked at ATM"}, {"role": "assistant", "content": "card_blocked"}]}
{"prompt": [{"role": "system", "content": "Classify the intent."}, {"role": "user", "content": "card blocked at ATM"}], "completion": [{"role": "assistant", "content": "card_blocked"}]}
```

The first is **conversational language modeling**; the second is **conversational prompt-completion**. With prompt-completion data, TRL computes loss on the completion only by default, which is what you want for classification, extraction and style tasks. For multi-turn "messages" data, TRL offers an assistant-only loss option, which requires the model's chat template to mark assistant turns; check the TRL docs for your model.

## The hyperparameters that matter

| Setting | What it does | Orientation (verify on your task) |
|---|---|---|
| Learning rate | Step size of weight updates | LoRA commonly around 1e-4 to 2e-4; full fine-tuning much lower (around 1e-5) |
| Epochs | Passes over the data | 1–3; more risks memorizing |
| Effective batch size | per-device batch × gradient accumulation × GPUs | 16–64 is common |
| Max length | Truncation length in tokens | Cover your longest real examples; longer costs memory |
| Scheduler / warmup | How the learning rate changes | Cosine or linear decay with a short warmup |
| Packing | Concatenating short examples to fill sequences | Speeds training on short examples; check that it suits your data |

Start from the library's defaults and a known-good recipe, change **one thing at a time**, and compare on your validation set.

## Reading the training curves

- **Training loss** should fall; **validation loss** should fall then flatten.
- If validation loss **rises** while training loss keeps falling, you are **overfitting**: stop earlier, reduce epochs or learning rate, or add data.
- If both stay high, the task may be unclear, the data inconsistent, or the learning rate too low.
- Loss is a proxy. Always run your **task evaluation** (accuracy, rubric scores) on checkpoints, because lower loss does not always mean better behavior.

## Hands-on: SFT with TRL and LoRA

```python
# train_sft.py  (pip install trl peft datasets transformers accelerate)
import os
from datasets import load_dataset
from peft import LoraConfig
from trl import SFTConfig, SFTTrainer

BASE = os.getenv("BASE_MODEL", "Qwen/Qwen3-1.7B")   # choose an open model whose license fits your use
data = load_dataset("json", data_files={"train": "train.jsonl", "validation": "val.jsonl"})  # prompt/completion format

args = SFTConfig(
    output_dir="out/intent-lora",
    num_train_epochs=2,
    per_device_train_batch_size=8,
    gradient_accumulation_steps=2,
    learning_rate=2e-4,
    lr_scheduler_type="cosine",
    max_length=1024,
    logging_steps=10,
    eval_strategy="steps", eval_steps=100,
    save_strategy="steps", save_steps=100,
    bf16=True,                 # on GPUs that support bfloat16
    report_to="none",          # or "wandb"/"mlflow" for experiment tracking
)
peft_config = LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05, target_modules="all-linear", task_type="CAUSAL_LM")

trainer = SFTTrainer(model=BASE, args=args, train_dataset=data["train"],
                     eval_dataset=data["validation"], peft_config=peft_config)
trainer.train()
trainer.save_model("out/intent-lora/final")   # saves the LoRA adapter, not a full model copy
```

Argument names evolve between TRL releases; check the SFTTrainer documentation for your installed version (`pip show trl`).

## Quick inference check

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

tok = AutoTokenizer.from_pretrained(BASE)
model = PeftModel.from_pretrained(AutoModelForCausalLM.from_pretrained(BASE, dtype=torch.bfloat16, device_map="auto"),
                                  "out/intent-lora/final")
msgs = [{"role": "system", "content": "Classify the intent."}, {"role": "user", "content": "paisay cut gaye lekin recharge nahi hua"}]
enc = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt", return_dict=True).to(model.device)
out = model.generate(**enc, max_new_tokens=8, do_sample=False)
print(tok.decode(out[0][enc["input_ids"].shape[1]:], skip_special_tokens=True))
```

Some reasoning models emit a "thinking" section by default; for classification, disable it via the chat template options described on the model card, or train the model to answer directly.

## Worked example: an e-commerce intent model in Karachi

An online marketplace fine-tunes a 1.7B model on 5,400 prompt-completion examples across 14 intents in Urdu, Roman Urdu and English. First run: validation loss rises after epoch 1 and accuracy on the validation set peaks at the epoch-1 checkpoint. They keep the epoch-1 adapter, add 600 examples for two confused intents, and retrain; accuracy on the untouched test set improves (numbers illustrative), and latency is a fraction of their previous API calls.

## Pitfalls

- Computing loss on the prompt as well as the answer for short-answer tasks.
- A different chat template at serving time (for example a runtime that does not apply the model's template).
- Choosing the final checkpoint by default rather than the best validation checkpoint.
- Tuning hyperparameters on the test set.

## How to measure success

Validation curves that show no overfitting at the chosen checkpoint, task metrics on the test set that beat your baseline from the customization ladder, and an adapter you can load and run with the production template.

## Video lecture: Supervised fine-tuning fundamentals

Lecture coming soon · 15 chapters · about 8 minutes. Read the full transcript below.

1. SFT fundamentals
2. Analogy: learning from a folder of great replies
3. What SFT does
4. Three make-or-break details
5. Data formats
6. Hyperparameters (orientation)
7. Reading curves
8. Hands-on: train_sft.py
9. Worked example: Karachi marketplace (illustrative)
10. Simple example: gym chain
11. Productive habits
12. FAQ: small fine-tune vs big prompted model?
13. Try this now
14. Watch me do it
15. Recap

## Lecture transcript

### SFT fundamentals

Supervised fine-tuning is the workhorse of model customization. It is also where most first attempts quietly go wrong, not because the idea is hard, but because of three small details. In this lesson you will see what actually happens during SFT, the data formats that make it easy, the hyperparameters that matter, how to read training curves, and a complete working script with TRL and LoRA.

### Analogy: learning from a folder of great replies

An analogy for SFT. Imagine teaching someone to write customer replies by giving them a folder of excellent past replies, each paired with the customer's message. They read message, reply, message, reply, thousands of times, and slowly their instincts shift. That is SFT. Loss masking means you only grade them on the replies they write, not on copying out the customer's message. And the chat template is the letterhead: if you train on one letterhead and then hand them a different one at work, they hesitate.

### What SFT does

A language model predicts the next token. SFT simply continues that training on your examples. For each example the model sees the prompt and the ideal response, and its weights are nudged so the ideal response becomes a little more likely. Repeat over thousands of examples, for one to a few passes called epochs, and the model's default behavior shifts toward your demonstrations.

### Three make-or-break details

Three details make or break it. First, loss masking. You want the model to learn to produce answers, not to reproduce your prompts, so compute the loss on the response tokens only. Second, the chat template. Instruction-tuned models expect conversations wrapped in special tokens, and training and serving must use the same template. Third, consistency. The system prompt, formatting and style in training must match production exactly.

### Data formats

TRL, Hugging Face's training library, understands conversational data. The simplest reliable shape for classification and extraction is prompt and completion: the prompt holds the system and user messages, the completion holds the assistant's answer. With that shape, TRL computes loss on the completion only by default. For multi-turn message data there is an assistant-only loss option, but it depends on the model's chat template marking assistant turns, so check the docs for your model.

### Hyperparameters (orientation)

The hyperparameters that matter most. Learning rate: for LoRA, something around one to two times ten to the minus four is a common starting point, and full fine-tuning uses much lower rates. Epochs: one to three. Effective batch size, which is batch per device times gradient accumulation times GPUs, often sixteen to sixty-four. Maximum length, enough to cover your real examples. And a scheduler with a short warmup. Start from defaults, change one thing at a time, and compare on validation.

### Reading curves

Now read the curves. Training loss should fall. Validation loss should fall and then flatten. If validation loss starts rising while training loss keeps falling, you are overfitting: the model is memorizing. Use an earlier checkpoint, fewer epochs, a lower learning rate or more data. If both stay high, suspect unclear labels or too low a learning rate. And remember that loss is a proxy. Always run your task evaluation on the checkpoints.

### Hands-on: train_sft.py

The script in the lesson is complete. It loads prompt-completion JSON lines for training and validation, sets up SFTConfig with two epochs, a learning rate of two times ten to the minus four, cosine decay, a maximum length, evaluation and saving every hundred steps. It adds a LoRA configuration on all linear layers, trains, and saves the adapter. A second snippet loads the base model plus adapter and classifies a Roman Urdu message. Argument names evolve, so check the docs for your TRL version.

### Worked example: Karachi marketplace (illustrative)

A worked example with illustrative numbers. A Karachi marketplace fine-tunes a one point seven billion model on fifty-four hundred examples across fourteen intents in Urdu, Roman Urdu and English. Validation loss rises after the first epoch, and validation accuracy peaks there too. They keep that checkpoint, add six hundred examples for two intents the model confused, and retrain. Test accuracy improves, and responses come back far faster than their previous API calls.

### Simple example: gym chain

A simple example. A gym chain wants a tiny model to label member messages as booking, billing, complaint or other. They prepare eight hundred prompt and completion examples, train a LoRA adapter for two epochs on a small instruct model, and watch the curves. Validation loss flattens after the first epoch. They compare checkpoints on the validation set: epoch one scores slightly better. They keep epoch one, run the test set once, and move on. The whole job takes less than an hour on one GPU.

### Productive habits

A few habits make SFT experiments far more productive. Start with a tiny run on a hundred examples to confirm the pipeline works end to end, including evaluation, before spending hours of GPU time. Keep a fixed validation set and never change it mid-project. Change one hyperparameter at a time. And save the exact config and data hash with every run, so when something works, you know precisely why.

### FAQ: small fine-tune vs big prompted model?

A frequent question: how do I know whether to fine-tune a smaller model or just use a bigger one with a good prompt? Run both on your golden set. If the small fine-tuned model matches the big prompted one on your task, you have won on cost and speed. If it falls short, check whether more or better data closes the gap. And remember maintenance: the small model is now yours to retrain and monitor, while the big model's prompt is simpler to update.

### Try this now

Try this now. Convert twenty of your labeled examples into the prompt and completion format shown in the lesson, with the exact system prompt you will use in production. Load them with the datasets library and print one example after the tokenizer's chat template is applied. Seeing the special tokens around your text is the moment chat templates stop being abstract.

### Watch me do it

Watch me do it. I convert two thousand labeled messages into prompt and completion lines with our exact production system prompt. I load them with the datasets library and print one example after the chat template is applied, to see the special tokens. Then I do a smoke run: one hundred examples, twenty steps, to confirm training, saving and evaluation all work. It does. Now the real run: two epochs, learning rate two times ten to the minus four, LoRA on all linear layers, evaluation every hundred steps. I watch the curves: training loss falls steadily; validation loss flattens after about one and a half epochs. I evaluate three saved checkpoints on the validation set with macro F1: the middle one wins. I load it with the base model, test five tricky Roman Urdu messages by hand, and only then run the test set, once. Result recorded with the config and data hash.

### Recap

Recap. SFT continues next-token training on your demonstrations. Mask loss to responses, keep templates and prompts identical to production, and use prompt-completion format for short-answer tasks. Tune one thing at a time, watch for overfitting, and pick checkpoints by task metrics. Your next step: train a LoRA adapter on your dataset, plot the curves, and compare accuracy across checkpoints.

## Key takeaways

- SFT continues next-token training on your demonstrations
- Mask loss to completions; keep chat template and system prompt identical to production
- Prompt-completion format gives completion-only loss by default in TRL
- Tune learning rate, epochs and batch size one at a time; watch validation loss for overfitting
- Pick checkpoints by task metrics on validation, confirm once on test

## Try it

Train a LoRA SFT adapter on your prepared dataset, plot training and validation loss, and compare task accuracy at each saved checkpoint.

- [Previous: Training data safety, licensing and privacy](https://optimizeall.com/learn/fine-tuning-and-custom-models/data-safety-licensing-and-privacy)
- [Next: LoRA, QLoRA and parameter-efficient fine-tuning](https://optimizeall.com/learn/fine-tuning-and-custom-models/lora-qlora-and-peft)
- [All lessons of Fine-Tuning, Distillation and Custom Models](https://optimizeall.com/learn/fine-tuning-and-custom-models)
