Fine-Tuning, Distillation and Custom ModelsSupervised fine-tuning and PEFT · Lesson 7 of 16

LoRA, QLoRA and parameter-efficient fine-tuning

Article · 16 min · 9 min lecture

Video lecture

LoRA, QLoRA and parameter-efficient fine-tuning

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

LoRA, QLoRA and PEFT

  • How they work
  • Choosing settings
  • From adapter to served model

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The problem PEFT solves

Full fine-tuning updates every weight. For a 7B model that means storing gradients and optimizer states for billions of parameters, typically needing several times the model's size in GPU memory, and producing a full new copy of the model per task. Parameter-efficient fine-tuning (PEFT) freezes the base model and trains a small number of new parameters instead. The most widely used method is LoRA.

How LoRA works

For a weight matrix W of size d × k, LoRA learns a low-rank update:

W' = W + (alpha / r) × B × A        where A is r × k, B is d × r, and r << min(d, k)
  • r (rank) controls capacity: typical values 8–64.
  • alpha scales the update; a common convention is alpha = r or 2r.
  • A and B are the only trained weights (the adapter); W stays frozen.

Parameter maths (illustrative): a 4096 × 4096 matrix has 16.8M weights. LoRA with r = 16 trains 16 × (4096 + 4096) = 131,072 parameters, about 0.8% of that matrix. Across a whole model, adapters are typically tens to a few hundred MB instead of many GB.

Which layers? Early practice applied LoRA only to attention projections. Current guidance commonly applies it to all linear layers (attention and MLP), which tends to work better. Research published in 2025 (for example "LoRA Without Regret" by Thinking Machines) reported that LoRA applied to all layers can match full fine-tuning on many post-training datasets, with a learning rate roughly an order of magnitude higher than full fine-tuning, while noting capacity limits for very large datasets.

QLoRA: fine-tuning on smaller GPUs

QLoRA loads the frozen base model in 4-bit (the NF4 "normal float" data type, with double quantization of the scales) and trains LoRA adapters in 16-bit on top. The original QLoRA paper showed fine-tuning a 65B model on a single 48 GB GPU. In practice, QLoRA lets you fine-tune 7–14B models on a single 16–24 GB consumer GPU, with somewhat slower training and a small quality trade-off compared with 16-bit LoRA.

Memory sketch (illustrative): 8B model in 4-bit ≈ 5 GB of weights, plus LoRA parameters and optimizer states (small), plus activations (depend on batch size and sequence length; gradient checkpointing reduces them). Many 8B QLoRA runs fit in 16–24 GB with modest batch sizes and sequence lengths.

Variants you will see

  • DoRA (weight-decomposed LoRA) and rsLoRA (rank-stabilized scaling) are refinements supported in the PEFT library; try them only after a solid LoRA baseline.
  • Adapters per task: keep one base model and many small adapters (one per client, language or task), then serve them together (multi-LoRA serving, Module 6).
  • Merging: fold an adapter into the base weights for simpler deployment (merge_and_unload() in PEFT), at the cost of losing hot-swappability.

Hands-on: QLoRA with Transformers, PEFT and TRL

# train_qlora.py  (pip install trl peft bitsandbytes transformers datasets accelerate)
import os, torch
from datasets import load_dataset
from peft import LoraConfig
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
from trl import SFTConfig, SFTTrainer

BASE = os.getenv("BASE_MODEL", "Qwen/Qwen3-8B")
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
                         bnb_4bit_use_double_quant=True, bnb_4bit_compute_dtype=torch.bfloat16)
model = AutoModelForCausalLM.from_pretrained(BASE, quantization_config=bnb, device_map="auto", dtype=torch.bfloat16)

lora = LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05, target_modules="all-linear", task_type="CAUSAL_LM")
args = SFTConfig(output_dir="out/voice-qlora", num_train_epochs=2, per_device_train_batch_size=4,
                 gradient_accumulation_steps=4, learning_rate=2e-4, gradient_checkpointing=True,
                 max_length=2048, bf16=True, logging_steps=10, eval_strategy="steps", eval_steps=50, report_to="none")
data = load_dataset("json", data_files={"train": "train.jsonl", "validation": "val.jsonl"})
trainer = SFTTrainer(model=model, args=args, train_dataset=data["train"], eval_dataset=data["validation"], peft_config=lora)
trainer.train()
trainer.save_model("out/voice-qlora/final")

Check the Transformers and bitsandbytes docs for your GPU; bitsandbytes 4-bit support depends on hardware and platform.

Merging for deployment

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base = AutoModelForCausalLM.from_pretrained(BASE, dtype=torch.bfloat16)      # load base in 16-bit to merge
merged = PeftModel.from_pretrained(base, "out/voice-qlora/final").merge_and_unload()
merged.save_pretrained("out/voice-merged", safe_serialization=True)
AutoTokenizer.from_pretrained(BASE).save_pretrained("out/voice-merged")
# Then quantize the merged model for serving (e.g. GGUF for llama.cpp/Ollama, or AWQ/FP8 for vLLM)

Merge into a 16-bit base, not the 4-bit training copy, then quantize the merged result for serving and re-evaluate after quantization.

Choosing r and target modules

SituationSuggested start
Style or format change, few thousand examplesr = 8–16, all-linear
Classification across many classes/languagesr = 16–32, all-linear
Large datasets or harder behaviorsr = 32–64; compare against full fine-tuning if feasible

Worked example: brand voice for a creator agency in London

An agency manages 12 creators, each with a distinct voice. They train one QLoRA adapter per creator on 400–900 approved captions each, on a single 24 GB GPU, overnight. Each adapter is under a few hundred MB. A shared base model serves all 12 via multi-LoRA, and editors blind-rate outputs: most adapters beat the prompted base model on "sounds like the creator" (illustrative). Two creators with fewer than 300 examples show little gain; the agency collects more before retraining.

Pitfalls

  • Merging into the quantized 4-bit weights instead of a 16-bit base.
  • Using too small a rank for a large, diverse dataset (underfitting) or too large for a tiny one (overfitting).
  • Forgetting to re-evaluate after merging and quantizing for serving.
  • Mismatched tokenizer or chat template between training and serving.

How to measure success

An adapter trained within your GPU budget that beats the prompted baseline on your evaluation, and a merged, quantized serving artifact whose scores are re-verified.

Key takeaways

  • LoRA trains small low-rank adapters (A, B) while the base stays frozen
  • Applying LoRA to all linear layers is the common current default
  • QLoRA loads the base in 4-bit NF4 so 7–14B models can be tuned on a single consumer GPU
  • Adapters enable one base model with many tasks; merging simplifies deployment
  • Merge into a 16-bit base, quantize for serving, and re-evaluate

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A 4096 × 4096 matrix gets a LoRA adapter with r = 8. How many trainable parameters does that adapter add?
  2. What does QLoRA change compared with LoRA?
  3. You trained with QLoRA and want a merged model for serving. What is the recommended approach?

Put it into practice

Train a QLoRA adapter for a style task on your GPU (or a rented one), merge it, quantize for serving, and compare scores before and after quantization.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.