AI · Advanced · 426 minutes · free · updated Sep 2026
Fine-Tuning, Distillation and Custom Models
When not to fine-tune, data curation, SFT, LoRA/QLoRA, DPO, GRPO, distillation, embeddings, evaluation, cost and deployment
- Lessons
- 16 in 6 modules
- Video lectures
- 16 lectures, 133 minutes
- Updated
- Sep 2026
Tools you'll use
- Hugging Face TRL
- PEFT
- Transformers
- Unsloth
- Axolotl
- bitsandbytes
- Sentence Transformers
- vLLM
- Vertex AI
- Amazon Bedrock
- Microsoft Foundry
- Presidio
- MLflow
- Ollama
About this course
Fine-tuning is powerful for the right problems and a costly detour for the wrong ones. This advanced, hands-on course teaches you to decide with evidence and then do it properly. You will climb the customization ladder before touching weights, map methods to gaps (SFT, LoRA and QLoRA, DPO and other preference methods, reinforcement fine-tuning with GRPO, distillation and embedding fine-tuning), and build datasets that teach what you mean: labeling guides, coverage, deduplication, leakage-free splits, synthetic data with strict filtering, PII redaction and license checks. You will train with Hugging Face TRL, PEFT, Unsloth and Axolotl, evaluate against your best non-fine-tuned baseline by slice with confidence intervals, compare hosted platforms including the 2026 changes, and model costs and break-even. Finally you will deploy, monitor drift and retrain, and complete a capstone: a fine-tuned support-intent router with a full evaluation report.
What you will learn
- Decide when fine-tuning is justified using the customization ladder
- Match gaps to SFT, LoRA/QLoRA, DPO, GRPO, distillation or embedding tuning
- Build clean, leakage-free, privacy-safe datasets, including filtered synthetic data
- Train adapters with TRL, PEFT, Unsloth or Axolotl reproducibly
- Evaluate fine-tunes against the best baseline by slice with confidence intervals
- Compare hosted and self-managed options and model costs and break-even
- Deploy, monitor drift and retrain custom models safely
Before you start
- Open-Weight and Local AI: Run, Choose and Deploy Your Own Models
- Evaluating and Monitoring LLM Applications
Course content
Decide before you tune
When fine-tuning is (and is not) the answer, the customization ladder, and a map of SFT, PEFT, preference, reinforcement, distillation and embedding methods.
- When NOT to fine-tune: the customization ladder — 15 min
- The customization methods map: SFT, PEFT, preference, RFT and distillation — 15 min
Training data: design, synthesis and safety
Design and curate datasets, generate and filter synthetic data, distil from teacher models, and handle privacy, licensing and safety.
- Dataset design and curation: quality beats quantity — 16 min
- Synthetic data generation and distillation — 16 min
- Training data safety, licensing and privacy — 15 min
Supervised fine-tuning and PEFT
How SFT works, LoRA and QLoRA in depth, merging adapters for deployment, and the open-source tooling that makes training fast and reproducible.
- Supervised fine-tuning fundamentals — 16 min
- LoRA, QLoRA and parameter-efficient fine-tuning — 16 min
- Open-source tooling: TRL, Unsloth, Axolotl and the training stack — 15 min
Preference and reinforcement fine-tuning
RLHF and DPO for subjective quality, and reinforcement fine-tuning with graders (GRPO) for checkable tasks, including reward hacking defenses.
- Preference optimization: RLHF and DPO for practitioners — 16 min
- Reinforcement fine-tuning with graders: RFT and GRPO — 16 min
Evaluation, hosted platforms, embeddings and cost
Prove a fine-tune beats your best baseline, choose hosted platforms with portability in mind, fine-tune embeddings for retrieval, and model costs and break-even.
- Evaluation before and after fine-tuning — 16 min
- Hosted fine-tuning: which platforms, which models, which trade-offs — 15 min
- Fine-tuning embedding models for better retrieval — 15 min
- Cost modeling: training, inference and break-even — 15 min
Deploy, monitor and capstone
Serve and version custom models, monitor drift and quality, plan retraining, and build a fine-tuned support-intent router with a full evaluation report.
- Deploying custom models and monitoring for drift — 16 min
- Capstone: fine-tune a small model for support-intent routing — 20 min
Certificate: Certified Custom Model Engineer
The holder can decide when fine-tuning is justified, build clean and privacy-safe datasets including filtered synthetic data, train LoRA and QLoRA adapters, apply DPO and GRPO, fine-tune embedding models, evaluate against the best baseline with slices and confidence intervals, model costs and break-even, and deploy and monitor custom models with drift-triggered retraining.
- Final assessment
- 25 questions, 40 minutes
- Passing score
- 80%