Skip to content

Fine-Tuning, Distillation and Custom Models · Supervised fine-tuning and PEFT · lesson 8 of 16 · 15 min

Open-source tooling: TRL, Unsloth, Axolotl and the training stack

The stack, layer by layer

| Layer | Common choices | Role | |---|---|---| | Models and tokenizers | Hugging Face Transformers | Load models, templates, generation | | Datasets | Hugging Face Datasets | Load, map and stream JSONL/Parquet | | PEFT | Hugging Face PEFT | LoRA, QLoRA, DoRA adapters | | Trainers | TRL (SFT, DPO, GRPO, reward modeling, distillation trainers) | Training loops for post-training methods | | Speed/memory optimizers | Unsloth | Faster, lower-memory fine-tuning via custom kernels; notebooks for many models | | Config-driven orchestration | Axolotl, LLaMA-Factory | YAML-configured training pipelines, multi-GPU, many methods | | Distributed training | Accelerate, DeepSpeed, FSDP | Multi-GPU and multi-node scaling | | Tracking | Weights & Biases, MLflow, TensorBoard | Metrics, configs, artifacts | | Evaluation | Your eval harness, lm-evaluation-harness, promptfoo, Inspect | Before/after tests |

You do not need all of them. A solid default for a single GPU: TRL + PEFT (maximum transparency) or Unsloth (maximum speed on limited VRAM). For teams running many experiments or multi-GPU jobs: Axolotl or LLaMA-Factory configs checked into Git.

TRL in one paragraph

TRL (Transformer Reinforcement Learning) is Hugging Face's post-training library. Its trainers share a pattern: a *Config object for hyperparameters, a *Trainer that accepts a model (name or object), datasets and an optional peft_config. You saw SFTTrainer; DPOTrainer and GRPOTrainer follow the same shape (Module 4). TRL also provides a CLI (trl sft ..., trl dpo ...) for quick runs. It moves fast; pin versions and read the release notes when upgrading.

Unsloth: speed and memory on one GPU

Unsloth patches model code with optimized kernels to fine-tune faster and with less memory, and integrates with TRL trainers.

# unsloth_sft.py  (install per Unsloth docs for your CUDA/GPU)
from unsloth import FastLanguageModel
from datasets import load_dataset
from trl import SFTConfig, SFTTrainer

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="unsloth/Qwen3-8B", max_seq_length=2048, load_in_4bit=True)   # check Unsloth's model list
model = FastLanguageModel.get_peft_model(
    model, r=16, lora_alpha=16, lora_dropout=0,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"],
    use_gradient_checkpointing="unsloth")

data = load_dataset("json", data_files="train.jsonl", split="train")
trainer = SFTTrainer(model=model, processing_class=tokenizer, train_dataset=data,
                     args=SFTConfig(output_dir="out/unsloth", per_device_train_batch_size=2,
                                    gradient_accumulation_steps=8, num_train_epochs=2,
                                    learning_rate=2e-4, logging_steps=10, report_to="none"))
trainer.train()
model.save_pretrained("out/unsloth/adapter"); tokenizer.save_pretrained("out/unsloth/adapter")
# Unsloth also offers helpers to export merged models and GGUF files; see its docs.

Axolotl: training as configuration

Axolotl expresses a run as YAML, which makes experiments reproducible and reviewable in pull requests.

# qlora-intents.yml  (check Axolotl docs for current keys)
base_model: Qwen/Qwen3-8B
load_in_4bit: true
adapter: qlora
lora_r: 16
lora_alpha: 32
lora_dropout: 0.05
lora_target_linear: true

datasets:
  - path: data/train.jsonl
    type: chat_template
val_set_size: 0.05
sequence_len: 2048
sample_packing: true

micro_batch_size: 4
gradient_accumulation_steps: 4
num_epochs: 2
learning_rate: 0.0002
lr_scheduler: cosine
bf16: auto
gradient_checkpointing: true

output_dir: ./out/qlora-intents
axolotl train qlora-intents.yml

Where to get GPUs

  • Your own workstation (see the open-weight course's hardware lesson) for frequent small jobs.
  • Cloud GPU instances from major clouds (choose a region that fits your data rules) or GPU-focused providers, billed hourly.
  • Notebook platforms for experiments.
  • Managed training on cloud ML platforms (Vertex AI, SageMaker, Azure ML) when you want job scheduling and governance.

For sensitive data, prefer infrastructure in an approved region with encryption and access controls; do not upload regulated data to free notebook services.

Reproducibility checklist

  • Pin library versions (requirements.txt or lock file) and record the base model revision (commit hash).
  • Fix random seeds; log the full config and dataset version (hash of the training files).
  • Track every run (loss curves, eval metrics, hardware, duration, cost).
  • Store adapters with a model card: base model, data card link, metrics, intended use, limitations.

Worked example: from notebook to pipeline

A Riyadh fintech's first fine-tune was a notebook on one engineer's laptop. Results were good, but nobody could reproduce them after a library update. They moved to Axolotl YAML in Git, pinned versions in a container image, ran training on cloud GPUs in an in-kingdom region, logged runs to MLflow, and added an evaluation step that blocks promotion if accuracy drops. The second model took less engineer time than the first, despite being larger.

Pitfalls

  • Mixing incompatible versions of Transformers, TRL, PEFT and bitsandbytes (pin them together).
  • No experiment tracking, so the "good run" cannot be found or reproduced.
  • Training on sensitive data in unapproved environments.
  • Copying a config for a different model family without checking chat template and target modules.

How to measure success

Any team member can reproduce a training run from Git (config, versions, data hash) and get matching metrics within normal variance.

Video lecture: Open-source tooling: TRL, Unsloth, Axolotl and the training stack

Lecture coming soon · 15 chapters · about 8 minutes. Read the full transcript below.

  1. The training stack
  2. Analogy: a professional kitchen
  3. The stack
  4. Choosing
  5. Unsloth
  6. Axolotl
  7. GPU sources
  8. Reproducibility checklist
  9. Worked example: Riyadh fintech
  10. Simple example: 12 GB GPU
  11. Track every run
  12. FAQ: own pipeline or managed?
  13. Try this now
  14. Watch me do it
  15. Recap

Lecture transcript

The training stack

Your first fine-tune usually lives in a notebook on one person's machine. It works, everyone celebrates, and three weeks later nobody can reproduce it. In this lesson you will map the open-source training stack, see when to use TRL, Unsloth or Axolotl, where to get GPUs, and how to make training reproducible so the second model is easier than the first.

Analogy: a professional kitchen

Think of the training stack like a professional kitchen. Transformers is the pantry with every ingredient. PEFT is a technique, like sous vide, that gets great results with less energy. TRL is the set of recipes: SFT, DPO, GRPO. Unsloth is a faster oven. Axolotl is the kitchen's written recipe cards that any chef can follow exactly. And experiment tracking is the kitchen logbook: what was cooked, when, how, and how it tasted.

The stack

Here is the stack, layer by layer. Transformers loads models and templates. Datasets handles your data. PEFT provides LoRA and QLoRA adapters. TRL provides the trainers, SFT, DPO, GRPO, reward modeling and distillation. Unsloth speeds up training and cuts memory with custom kernels. Axolotl and LLaMA-Factory turn runs into YAML configuration. Accelerate, DeepSpeed and FSDP scale to multiple GPUs. And tracking tools like Weights and Biases or MLflow record every run.

Choosing

You do not need all of it. On a single GPU, TRL plus PEFT gives maximum transparency, and Unsloth gives maximum speed when memory is tight. For teams running many experiments or multi-GPU jobs, a configuration tool like Axolotl, checked into Git, keeps everyone honest. TRL itself follows one pattern across trainers: a config object, a trainer that takes a model, datasets and an optional adapter config. It moves fast, so pin versions and read release notes.

Unsloth

Unsloth in practice. You load a model with its fast language model helper, in four-bit if you like, then add LoRA adapters with get peft model, choosing rank, alpha and target modules, and gradient checkpointing. You hand that model to TRL's SFT trainer as usual. Unsloth also offers helpers to export merged models and GGUF files. Check its model list and installation guide for your GPU.

Axolotl

Axolotl expresses a run as YAML. Base model, four-bit loading, QLoRA adapter, rank and alpha, all linear targets, the dataset path with a chat template type, validation split, sequence length, packing, batch sizes, epochs, learning rate, scheduler and output directory. One command, axolotl train, runs it. The magic is not the command, it is that the whole experiment is a reviewable file in a pull request.

GPU sources

Where do the GPUs come from? Your own workstation for frequent small jobs. Hourly cloud GPU instances, in a region that fits your data rules. Notebook platforms for experiments. Or managed training on platforms like Vertex AI, SageMaker or Azure Machine Learning when you want scheduling and governance. One firm rule: regulated customer data never goes into a free public notebook. Use approved infrastructure with encryption and access control.

Reproducibility checklist

Now reproducibility, the habit that pays for itself. Pin library versions together, because Transformers, TRL, PEFT and bits and bytes must match. Record the base model's exact revision. Fix random seeds. Hash your training files so you know which data made which model. Track every run, with curves, metrics, hardware, duration and cost. And store each adapter with a model card describing base model, data, metrics, intended use and limitations.

Worked example: Riyadh fintech

A worked example. A Riyadh fintech's first fine-tune was a notebook on one laptop. Good results, but no one could reproduce them after a library update. For the second model, they moved to Axolotl YAML in Git, pinned versions in a container image, trained on cloud GPUs in an in-kingdom region, logged runs to MLflow, and added an evaluation gate that blocks promotion if accuracy drops. The second model was larger, yet took less engineering time.

Simple example: 12 GB GPU

A simple example. A two-person startup has one twelve gigabyte GPU. They try TRL and PEFT with QLoRA on a seven billion model and run out of memory at their sequence length. They switch the same data to Unsloth's loader with four-bit and gradient checkpointing, and the job fits and runs faster. Same data, same method, different tool. Knowing the stack lets you pick the tool that fits your constraints.

Track every run

Experiment tracking deserves its own moment. Every run should record its config, library versions, data hash, hardware, duration, cost, training curves and evaluation metrics, in one place the whole team can see. Tools like Weights and Biases or MLflow do this with a few lines of setup. It turns the question which run was the good one from an argument into a lookup, and it is the first thing an auditor or a new team member will ask for.

FAQ: own pipeline or managed?

A common question: should we build our own training pipeline or use a managed service? If you fine-tune rarely and want minimal operations, managed training on a cloud platform or an inference provider that supports adapters can be simpler. If you fine-tune often, need specific methods, or must keep data in particular environments, owning a small reproducible pipeline with open-source tools pays off. Many teams do both: prototype on a managed service, then move steady workloads to their own pipeline.

Try this now

Try this now. Create a requirements file that pins exact versions of transformers, TRL, PEFT, datasets and accelerate, and commit it next to your training script. Then write the base model's revision hash into your config. Those two small steps make your next run reproducible.

Watch me do it

Watch me do it. I turn a working notebook into a reproducible project. First I create a requirements file pinning transformers, TRL, PEFT, datasets, accelerate and bits and bytes to the exact versions that worked. Then I write the training settings as an Axolotl YAML: base model with its revision hash, QLoRA, rank, dataset path with the chat template type, sequence length, batch sizes, epochs and learning rate. I compute the SHA of the training file and write it into the config comments. I commit everything to Git. Next I run axolotl train on a rented cloud GPU in our approved region, with runs logged to MLflow. The run finishes; I note the metrics. Then a colleague clones the repository on a different machine and runs the same command. Her validation macro F1 matches mine to the second decimal place. That is the moment the model stops belonging to one person's laptop.

Recap

Recap. Use TRL and PEFT for transparency, Unsloth for speed on one GPU, and Axolotl or LLaMA-Factory for reproducible configuration. Get GPUs from environments that match your data rules. Pin, hash, track and document every run. Your next step: turn your SFT script into a pinned project or an Axolotl YAML in Git, run it twice, and confirm the metrics match within normal variance.

Key takeaways

  • TRL + PEFT is the transparent default; Unsloth maximizes speed and memory on one GPU
  • Axolotl and LLaMA-Factory make training reproducible YAML configuration
  • Accelerate, DeepSpeed and FSDP scale to multiple GPUs
  • Pin versions, record base model revision and data hash, and track every run
  • Use approved environments and regions for sensitive training data

Try it

Convert your SFT script into either an Axolotl YAML or a pinned TRL project in Git, run it twice, and confirm the metrics match within variance.