Skip to content

Fine-Tuning, Distillation and Custom Models · Preference and reinforcement fine-tuning · lesson 9 of 16 · 16 min

Preference optimization: RLHF and DPO for practitioners

Why preferences, not demonstrations

SFT needs you to write the ideal answer. For many qualities (tone, helpfulness, concision, safety, "sounds like us") humans are much better at comparing two answers than writing a perfect one. Preference optimization learns from those comparisons: given a prompt, a chosen response and a rejected response, shift the model toward outputs like the chosen one.

RLHF in one diagram

The classic pipeline (popularized by InstructGPT in 2022) has three stages:

1. SFT model  ──►  2. Reward model trained on human comparisons  ──►  3. RL (e.g. PPO) optimizes the SFT model
                                                                       to maximize reward, with a KL penalty
                                                                       keeping it close to the SFT model

It works, and frontier labs still use RL heavily in post-training, but classic RLHF is complex to run: two models to train, unstable RL, reward hacking (the policy exploits flaws in the reward model), and significant compute.

DPO: preferences without a reward model

Direct Preference Optimization (DPO) (Rafailov et al., 2023) showed that the RLHF objective can be optimized directly on preference pairs with a simple classification-style loss. Intuition:

  • Keep a frozen reference model (usually your SFT model).
  • For each pair, increase how much more likely the policy makes the chosen response relative to the reference, and decrease the same for the rejected response.
  • A parameter beta controls how far the model may drift from the reference: higher beta keeps it closer; lower beta allows bigger changes (and bigger risks).

No separate reward model, no RL sampling loop, and it trains like SFT. That is why DPO and its variants became the practitioner's default for preference tuning.

The family of variants (know the names)

| Method | Key idea | When to consider | |---|---|---| | DPO | Pairwise preferences + reference model | Default choice | | IPO | Regularized variant addressing DPO overfitting on deterministic preferences | If DPO overfits | | KTO | Learns from unpaired thumbs-up / thumbs-down labels | You have ratings, not pairs | | ORPO | Combines SFT and preference in one stage, no reference model | Simpler pipelines, smaller compute | | SimPO | Reference-free, length-normalized reward | Reducing length bias | | Online / iterative DPO | Generate new pairs from the current model and repeat | Continued improvement |

TRL implements DPO and several variants; check its docs for current options.

Where preference data comes from

  • Agent edits: the model's draft (rejected) vs the agent's edited version (chosen). Free, high-signal data from normal work.
  • A/B choices: show two drafts, let the user or reviewer pick.
  • Thumbs up/down: unpaired signals (use KTO or build pairs from same-prompt samples).
  • AI feedback: a strong judge model picks between samples against a rubric (RLAIF). Cheaper, but calibrate against humans.

Quality rules: pairs must differ on the dimension you care about; avoid pairs where "chosen" is merely longer (models learn length bias quickly); balance topics and languages; keep a held-out preference set to measure win rate.

Hands-on: DPO with TRL on top of your SFT model

{"prompt": [{"role": "user", "content": "My order is late again. Unacceptable."}], "chosen": [{"role": "assistant", "content": "I'm sorry, that's frustrating. I've checked order 4417: it's with the courier and due tomorrow. I've added a delivery-fee refund to your account."}], "rejected": [{"role": "assistant", "content": "Delays happen sometimes. Please wait."}]}
# train_dpo.py  (pip install trl peft datasets transformers accelerate)
import os
from datasets import load_dataset
from peft import LoraConfig
from trl import DPOConfig, DPOTrainer

SFT_MODEL = os.getenv("SFT_MODEL", "out/voice-merged")      # start from your merged SFT model
data = load_dataset("json", data_files={"train": "prefs_train.jsonl", "validation": "prefs_val.jsonl"})

args = DPOConfig(
    output_dir="out/voice-dpo",
    beta=0.1,                        # higher = stay closer to the reference model
    learning_rate=5e-6,              # preference tuning typically uses a lower LR than SFT; tune on validation
    num_train_epochs=1,
    per_device_train_batch_size=2,
    gradient_accumulation_steps=8,
    max_length=1024,
    bf16=True,
    eval_strategy="steps", eval_steps=50,
    report_to="none",
)
lora = LoraConfig(r=16, lora_alpha=32, target_modules="all-linear", task_type="CAUSAL_LM")
# With a peft_config, TRL can use the base weights (adapter disabled) as the implicit reference model.
trainer = DPOTrainer(model=SFT_MODEL, args=args, train_dataset=data["train"],
                     eval_dataset=data["validation"], peft_config=lora)
trainer.train()
trainer.save_model("out/voice-dpo/final")

Watch the logged preference metrics: reward accuracy (how often chosen scores above rejected) and margins should rise on validation. But the real test is a blind win-rate evaluation: show reviewers the SFT output and the DPO output for held-out prompts, in random order, and count wins.

Worked example: tone for a UAE bank's chat assistant

The bank's SFT model was accurate but curt. Over four weeks, agents' edits to drafts produced 3,200 pairs (draft = rejected, edited = chosen) in English and Arabic. The team filtered pairs where edits only changed length, balanced topics, and ran DPO with LoRA. In a blind review of 200 held-out conversations, reviewers preferred the DPO model in a clear majority of cases, with no drop on the factual-accuracy checklist (illustrative). Response length rose slightly; they monitored it to avoid verbosity creep.

Pitfalls

  • Running DPO without SFT first (the model cannot yet do the task).
  • Length bias: "chosen" answers systematically longer teaches verbosity.
  • Too low beta or too high learning rate: the model drifts, and quality on other tasks drops.
  • Measuring only training metrics, not blind human win rate and regression tests.

How to measure success

A blind win rate over the SFT model on held-out prompts, stable scores on accuracy and safety regression tests, and no unwanted shift in length or refusals.

Video lecture: Preference optimization: RLHF and DPO for practitioners

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

  1. Preference optimization
  2. Analogy: tasting panel
  3. Classic RLHF
  4. DPO
  5. Variants
  6. Preference data sources
  7. Hands-on: train_dpo.py
  8. Measuring success
  9. Worked example: UAE bank (illustrative)
  10. Simple example: hotel assistant (illustrative)
  11. Continuous preference loop
  12. FAQ: how many pairs?
  13. Try this now
  14. Watch me do it
  15. Recap

Lecture transcript

Preference optimization

Try this. Write the perfect reply to an angry customer. Hard, right? Now I show you two replies and ask which is better. Easy. That gap, between writing and choosing, is why preference optimization exists. In this lesson you will understand RLHF at a practitioner level, see why DPO became the default, meet the variants, learn where preference data comes from, and run a DPO job with TRL.

Analogy: tasting panel

Here is the analogy. Supervised fine-tuning is like a cooking class where the teacher shows you the perfect dish. Preference optimization is like a tasting panel: they taste two versions and say this one is better. You never see a perfect recipe, but after hundreds of tastings, your cooking drifts toward what the panel prefers. Beta is how far you are allowed to drift from your original style in one go. Too far, too fast, and you may forget how to cook the basics.

Classic RLHF

Classic RLHF has three stages. First, supervised fine-tuning. Second, train a reward model on human comparisons, so it predicts which answer people prefer. Third, use reinforcement learning, often an algorithm called PPO, to push the model toward higher reward, with a penalty that keeps it close to the original. It works, and frontier labs still rely heavily on reinforcement learning. But it is complex: two models, unstable training, and reward hacking, where the model exploits flaws in the reward model.

DPO

DPO, direct preference optimization, published in twenty twenty-three, showed you can skip the reward model and the reinforcement loop. You keep a frozen reference model, usually your SFT model. For each pair, training makes the chosen response more likely relative to the reference, and the rejected response less likely. A parameter called beta controls how far the model may drift. Higher beta stays closer, lower beta allows bigger changes and bigger risks. It trains much like SFT, which is why practitioners adopted it so quickly.

Variants

You will hear other names. IPO regularizes DPO when it overfits. KTO learns from unpaired thumbs up and thumbs down, when you have ratings but not pairs. ORPO folds SFT and preferences into one stage without a reference model. SimPO is reference-free and normalizes for length. And online or iterative DPO generates fresh pairs from the current model and repeats. Start with DPO, and reach for a variant when you have a specific reason.

Preference data sources

Where does preference data come from? The richest source is often free: agent edits. The model's draft is the rejected answer, the agent's edited version is the chosen one. You can also run A/B choices, collect thumbs up and down, or use a strong judge model against a rubric, which is cheaper but must be calibrated against humans. Two quality rules. Pairs must differ on what you care about, and chosen answers must not simply be longer, or your model learns to ramble.

Hands-on: train_dpo.py

The hands-on starts from your merged SFT model. Your data file has a prompt, a chosen reply and a rejected reply per line. The DPO config sets beta to point one, a learning rate lower than SFT, one epoch, a modest batch with gradient accumulation, and evaluation steps. With a LoRA config, TRL can use the base weights with the adapter switched off as the reference model, which saves memory. Train, save, and then evaluate properly.

Measuring success

How do you know it worked? The logs show reward accuracy, how often chosen beats rejected, and margins. Both should rise on validation. But the real test is a blind win rate. Take held-out prompts, generate answers from the SFT model and the DPO model, shuffle them, and have reviewers pick the better one without knowing which is which. Also re-run your accuracy and safety regression tests, and watch response length.

Worked example: UAE bank (illustrative)

A worked example, with illustrative results. A UAE bank's SFT model was accurate but curt. In four weeks, agent edits produced thirty-two hundred pairs in English and Arabic. The team removed pairs where the only change was length, balanced topics, and ran DPO with LoRA. In a blind review of two hundred held-out conversations, reviewers clearly preferred the DPO model, with no drop on the factual checklist. Length rose slightly, so they kept monitoring it.

Simple example: hotel assistant (illustrative)

A simple example. A hotel's booking assistant gives correct but stiff answers. Over two weeks, the front-desk team is shown two reply drafts for real guest questions and taps the better one. They collect five hundred pairs. After removing pairs where the only difference was length, they run a small DPO job on the SFT model. In a blind test on fifty new questions, staff prefer the new replies most of the time, and booking accuracy stays the same.

Continuous preference loop

A practical workflow that works well. Ship the SFT model with a light review loop, where agents edit drafts before sending. Log the original draft and the edited version automatically. Every few weeks, filter those pairs, remove length-only edits, balance topics, and run a DPO update. Evaluate with a blind win rate and your regression suite, and release through champion and challenger. Your product improves continuously from normal work, without a separate labeling project.

FAQ: how many pairs?

A question people ask: how many preference pairs do we need? It depends on how different your target behavior is from the current model. For a focused tone adjustment, a few hundred high-quality pairs can move the needle. For broader changes, thousands. As always, quality beats quantity: pairs where the chosen answer is clearly better on the dimension you care about teach far more than pairs where the difference is subtle or just length. Start small, measure the blind win rate, and scale if it helps.

Try this now

Try this now. Take ten outputs from your current model and, for each, write a better version yourself or ask a colleague to. You now have ten preference pairs. Look at them together: is the better version longer every time? If so, you have found length bias before it reaches training.

Watch me do it

Watch me do it. Our logs hold four weeks of agent edits: eighteen hundred pairs of original draft and edited reply. I load them and compute the length difference for each pair. Four hundred pairs differ only by added sentences with the same meaning; I drop them to avoid teaching length. I balance topics so refunds do not dominate, and split off two hundred pairs for validation. I run the DPO script on our merged SFT model with beta zero point one and one epoch. In the logs, reward accuracy on validation climbs from about fifty percent to the high seventies, and margins widen. Now the real test: I generate replies from the SFT and DPO models for sixty held-out prompts, shuffle them, and ask three agents to pick the better one blind. The DPO model wins most comparisons. I rerun the accuracy checklist and the safety suite: no regressions. Average length rose a little; I set an alert on it.

Recap

Recap. Preference optimization learns from comparisons. RLHF uses a reward model and reinforcement learning, while DPO optimizes preferences directly with a reference model and a beta setting. Agent edits are gold, but watch length bias. And judge success by blind win rate plus regression tests. Your next step: collect two hundred pairs from your workflow, filter them, run a small DPO job, and measure the win rate on fifty held-out prompts.

Key takeaways

  • Preference optimization learns from comparisons (chosen vs rejected), easier than writing ideal answers
  • Classic RLHF: SFT → reward model → RL with KL penalty; powerful but complex
  • DPO optimizes preferences directly with a reference model and a beta parameter
  • Agent edits are a rich source of preference pairs; watch length bias
  • Judge success by blind win rate plus regression tests

Try it

Collect 200 preference pairs from edits or A/B choices in your workflow, filter for length bias, and run a small DPO job; measure blind win rate on 50 held-out prompts.