---
title: "LoRA, QLoRA and parameter-efficient fine-tuning"
description: "The problem PEFT solves Full fine-tuning updates every weight. For a 7B model that means storing gradients and optimizer states for billions of…"
url: https://optimizeall.com/learn/fine-tuning-and-custom-models/lora-qlora-and-peft
updated: 2026-10-05
---

Fine-Tuning, Distillation and Custom Models · Supervised fine-tuning and PEFT · lesson 7 of 16 · 16 min

# LoRA, QLoRA and parameter-efficient fine-tuning

## The problem PEFT solves

Full fine-tuning updates every weight. For a 7B model that means storing gradients and optimizer states for billions of parameters, typically needing several times the model's size in GPU memory, and producing a full new copy of the model per task. **Parameter-efficient fine-tuning (PEFT)** freezes the base model and trains a small number of new parameters instead. The most widely used method is **LoRA**.

## How LoRA works

For a weight matrix `W` of size d × k, LoRA learns a low-rank update:

```text
W' = W + (alpha / r) × B × A        where A is r × k, B is d × r, and r << min(d, k)
```

- `r` (**rank**) controls capacity: typical values 8–64.
- `alpha` scales the update; a common convention is alpha = r or 2r.
- `A` and `B` are the only trained weights (the **adapter**); `W` stays frozen.

**Parameter maths (illustrative):** a 4096 × 4096 matrix has 16.8M weights. LoRA with r = 16 trains 16 × (4096 + 4096) = 131,072 parameters, about 0.8% of that matrix. Across a whole model, adapters are typically tens to a few hundred MB instead of many GB.

**Which layers?** Early practice applied LoRA only to attention projections. Current guidance commonly applies it to **all linear layers** (attention and MLP), which tends to work better. Research published in 2025 (for example "LoRA Without Regret" by Thinking Machines) reported that LoRA applied to all layers can match full fine-tuning on many post-training datasets, with a learning rate roughly an order of magnitude higher than full fine-tuning, while noting capacity limits for very large datasets.

## QLoRA: fine-tuning on smaller GPUs

**QLoRA** loads the frozen base model in **4-bit** (the NF4 "normal float" data type, with double quantization of the scales) and trains LoRA adapters in 16-bit on top. The original QLoRA paper showed fine-tuning a 65B model on a single 48 GB GPU. In practice, QLoRA lets you fine-tune 7–14B models on a single 16–24 GB consumer GPU, with somewhat slower training and a small quality trade-off compared with 16-bit LoRA.

**Memory sketch (illustrative):** 8B model in 4-bit ≈ 5 GB of weights, plus LoRA parameters and optimizer states (small), plus activations (depend on batch size and sequence length; gradient checkpointing reduces them). Many 8B QLoRA runs fit in 16–24 GB with modest batch sizes and sequence lengths.

## Variants you will see

- **DoRA** (weight-decomposed LoRA) and **rsLoRA** (rank-stabilized scaling) are refinements supported in the PEFT library; try them only after a solid LoRA baseline.
- **Adapters per task:** keep one base model and many small adapters (one per client, language or task), then serve them together (multi-LoRA serving, Module 6).
- **Merging:** fold an adapter into the base weights for simpler deployment (`merge_and_unload()` in PEFT), at the cost of losing hot-swappability.

## Hands-on: QLoRA with Transformers, PEFT and TRL

```python
# train_qlora.py  (pip install trl peft bitsandbytes transformers datasets accelerate)
import os, torch
from datasets import load_dataset
from peft import LoraConfig
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
from trl import SFTConfig, SFTTrainer

BASE = os.getenv("BASE_MODEL", "Qwen/Qwen3-8B")
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
                         bnb_4bit_use_double_quant=True, bnb_4bit_compute_dtype=torch.bfloat16)
model = AutoModelForCausalLM.from_pretrained(BASE, quantization_config=bnb, device_map="auto", dtype=torch.bfloat16)

lora = LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05, target_modules="all-linear", task_type="CAUSAL_LM")
args = SFTConfig(output_dir="out/voice-qlora", num_train_epochs=2, per_device_train_batch_size=4,
                 gradient_accumulation_steps=4, learning_rate=2e-4, gradient_checkpointing=True,
                 max_length=2048, bf16=True, logging_steps=10, eval_strategy="steps", eval_steps=50, report_to="none")
data = load_dataset("json", data_files={"train": "train.jsonl", "validation": "val.jsonl"})
trainer = SFTTrainer(model=model, args=args, train_dataset=data["train"], eval_dataset=data["validation"], peft_config=lora)
trainer.train()
trainer.save_model("out/voice-qlora/final")
```

Check the Transformers and bitsandbytes docs for your GPU; bitsandbytes 4-bit support depends on hardware and platform.

## Merging for deployment

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base = AutoModelForCausalLM.from_pretrained(BASE, dtype=torch.bfloat16)      # load base in 16-bit to merge
merged = PeftModel.from_pretrained(base, "out/voice-qlora/final").merge_and_unload()
merged.save_pretrained("out/voice-merged", safe_serialization=True)
AutoTokenizer.from_pretrained(BASE).save_pretrained("out/voice-merged")
# Then quantize the merged model for serving (e.g. GGUF for llama.cpp/Ollama, or AWQ/FP8 for vLLM)
```

Merge into a 16-bit base, not the 4-bit training copy, then quantize the merged result for serving and **re-evaluate after quantization**.

## Choosing r and target modules

| Situation | Suggested start |
|---|---|
| Style or format change, few thousand examples | r = 8–16, all-linear |
| Classification across many classes/languages | r = 16–32, all-linear |
| Large datasets or harder behaviors | r = 32–64; compare against full fine-tuning if feasible |

## Worked example: brand voice for a creator agency in London

An agency manages 12 creators, each with a distinct voice. They train one QLoRA adapter per creator on 400–900 approved captions each, on a single 24 GB GPU, overnight. Each adapter is under a few hundred MB. A shared base model serves all 12 via multi-LoRA, and editors blind-rate outputs: most adapters beat the prompted base model on "sounds like the creator" (illustrative). Two creators with fewer than 300 examples show little gain; the agency collects more before retraining.

## Pitfalls

- Merging into the quantized 4-bit weights instead of a 16-bit base.
- Using too small a rank for a large, diverse dataset (underfitting) or too large for a tiny one (overfitting).
- Forgetting to re-evaluate after merging and quantizing for serving.
- Mismatched tokenizer or chat template between training and serving.

## How to measure success

An adapter trained within your GPU budget that beats the prompted baseline on your evaluation, and a merged, quantized serving artifact whose scores are re-verified.

## Video lecture: LoRA, QLoRA and parameter-efficient fine-tuning

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. LoRA, QLoRA and PEFT
2. Analogy: sticky notes on a textbook
3. Why PEFT
4. How LoRA works
5. Target modules
6. QLoRA
7. Hands-on
8. Starting points
9. Worked example: London creator agency (illustrative)
10. Simple example: 3 client voices
11. Where the memory goes
12. FAQ: does LoRA cause forgetting?
13. Try this now
14. Watch me do it
15. Recap

## Lecture transcript

### LoRA, QLoRA and PEFT

A few years ago, fine-tuning a seven billion parameter model meant a rack of expensive GPUs and a full copy of the model for every task. Today you can do it overnight on one consumer graphics card and store the result in a file smaller than a movie. That shift is thanks to LoRA and QLoRA. In this lesson you will learn how they work, how to choose settings, and how to take an adapter all the way to a served model.

### Analogy: sticky notes on a textbook

Here is the analogy for LoRA. Imagine a huge, expensive textbook. Instead of reprinting it with your changes, you add a slim set of sticky notes that adjust key pages. Anyone with the original book plus your sticky notes gets your version. You can keep different sets of notes for different courses, and swap them in seconds. QLoRA goes one step further: you keep a compressed pocket edition of the textbook on your desk while writing the notes, so it takes far less space.

### Why PEFT

Full fine-tuning updates every weight, which needs memory for gradients and optimizer states several times the model's size, and produces a full new model per task. Parameter-efficient fine-tuning freezes the base model and trains a small number of new parameters instead. LoRA is the most widely used way to do that.

### How LoRA works

Here is the idea. For each big weight matrix, LoRA learns a low-rank update made of two thin matrices, A and B, multiplied together and added to the frozen weights, scaled by alpha over r. The rank, r, controls capacity, usually eight to sixty-four. The numbers make it concrete. A four thousand by four thousand matrix has about sixteen point eight million weights. LoRA with rank sixteen trains about a hundred and thirty-one thousand. Under one percent. Across a whole model, adapters are tens to a few hundred megabytes.

### Target modules

Which layers get adapters? Early practice targeted only the attention projections. Current practice commonly targets all linear layers, attention and the feed-forward blocks, and that tends to work better. Research published in twenty twenty-five, for example the "LoRA Without Regret" study, reported that LoRA on all layers can match full fine-tuning on many post-training datasets, with a learning rate around ten times higher than full fine-tuning, while noting limits on very large datasets.

### QLoRA

QLoRA takes it further. It loads the frozen base model in four-bit precision, using a data type called NF4 with double quantization of the scales, and trains sixteen-bit LoRA adapters on top. The original paper fine-tuned a sixty-five billion parameter model on a single forty-eight gigabyte GPU. In practice, it lets you tune seven to fourteen billion parameter models on one sixteen to twenty-four gigabyte card, with slightly slower training and a small quality trade-off.

### Hands-on

The QLoRA script in the lesson loads the base with a bits and bytes four-bit configuration, adds a LoRA config on all linear layers, turns on gradient checkpointing to save activation memory, and trains with TRL's SFT trainer. For deployment, you load the base in sixteen-bit, attach the adapter, merge it with merge and unload, save, and then quantize the merged model for your serving engine. And then you re-evaluate, because quantization can shift behavior.

### Starting points

How to choose settings. For a style or format change with a few thousand examples, start with rank eight to sixteen on all linear layers. For classification across many classes and languages, sixteen to thirty-two. For large datasets or harder behaviors, thirty-two to sixty-four, and compare against full fine-tuning if you can. Refinements like DoRA and rank-stabilized LoRA exist, but try them only after a solid baseline.

### Worked example: London creator agency (illustrative)

A worked example, with illustrative results. A London agency manages twelve creators with distinct voices. They train one QLoRA adapter per creator on four to nine hundred approved captions each, on one twenty-four gigabyte GPU, overnight. A single base model serves all twelve adapters. In blind ratings, most adapters beat the prompted base model on sounding like the creator. Two creators with fewer than three hundred examples show little gain, so the agency collects more before retraining.

### Simple example: 3 client voices

A simple example. A marketing agency wants one small model to write product descriptions for three clients with very different tones: a luxury perfume brand, a budget electronics store and a children's toy shop. Instead of three full models, they train three LoRA adapters, each on about five hundred approved descriptions, on one twenty-four gigabyte GPU. At serving time, one base model loads, and the right adapter is picked per client. Three voices, one GPU, small files.

### Where the memory goes

What about memory for full fine-tuning versus LoRA, in practical terms? Full fine-tuning with a standard optimizer typically needs memory for the weights, the gradients and two optimizer states per parameter, several times the model size. LoRA keeps the base frozen and only needs gradients and optimizer states for the tiny adapter. QLoRA then shrinks the frozen base itself. That is the whole reason a job that once needed a multi-GPU server now fits on one card.

### FAQ: does LoRA cause forgetting?

A question I get: does LoRA make the model forget what it already knew? Much less than full fine-tuning, because the base weights stay frozen. But the adapter can still shift behavior in ways that hurt other tasks, especially with a high learning rate or many epochs. That is why your regression suite matters: after training, check general instructions, other languages and safety behavior, not only the target task. If you see forgetting, reduce epochs or learning rate, or mix in some general examples.

### Try this now

Try this now. Take a model you might fine-tune and look up two numbers in its config: hidden size and number of layers. Estimate the LoRA parameters for one projection matrix at rank sixteen: sixteen times the two dimensions added together. Then multiply by the number of target matrices per layer and the number of layers. Compare that with the model's total parameters. Seeing the percentage makes LoRA's efficiency real.

### Watch me do it

Watch me do it. I have a twenty-four gigabyte GPU and want to adapt an eight billion model to a client's brand voice. I load the base in four-bit with the bits and bytes config and check memory: about six gigabytes. I add a LoRA config, rank sixteen, alpha thirty-two, all linear layers, and print the trainable parameters: well under one percent of the model. I train on six hundred approved captions with gradient checkpointing; memory peaks around seventeen gigabytes. Training takes about forty minutes. Then deployment: I load the base in sixteen-bit on a bigger machine, attach the adapter, merge and save. I convert the merged model to GGUF, quantize to Q five K M, and load it in Ollama. Last and most important, I rerun the voice evaluation on the quantized model: the score is within noise of the unquantized adapter. Only now is it ready to serve.

### Recap

Recap. LoRA trains small adapters while the base stays frozen, and all linear layers is the usual default. QLoRA puts the base in four-bit so you can train on one consumer GPU. Keep adapters per task, or merge into a sixteen-bit base for simple deployment, then quantize and re-evaluate. Your next step: train a QLoRA adapter for a style task, merge and quantize it, and compare scores before and after quantization.

## Key takeaways

- LoRA trains small low-rank adapters (A, B) while the base stays frozen
- Applying LoRA to all linear layers is the common current default
- QLoRA loads the base in 4-bit NF4 so 7–14B models can be tuned on a single consumer GPU
- Adapters enable one base model with many tasks; merging simplifies deployment
- Merge into a 16-bit base, quantize for serving, and re-evaluate

## Try it

Train a QLoRA adapter for a style task on your GPU (or a rented one), merge it, quantize for serving, and compare scores before and after quantization.

- [Previous: Supervised fine-tuning fundamentals](https://optimizeall.com/learn/fine-tuning-and-custom-models/sft-fundamentals)
- [Next: Open-source tooling: TRL, Unsloth, Axolotl and the training stack](https://optimizeall.com/learn/fine-tuning-and-custom-models/open-source-training-tooling)
- [All lessons of Fine-Tuning, Distillation and Custom Models](https://optimizeall.com/learn/fine-tuning-and-custom-models)
