Fine-Tuning, Distillation and Custom ModelsDeploy, monitor and capstone · Lesson 16 of 16
Capstone: fine-tune a small model for support-intent routing
Video lecture
Capstone: fine-tune a small model for support-intent routing
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Capstone: support-intent router
This is where everything comes together. You are going to fine-tune a small open-weight model to route customer support messages to the right intent, fast and cheaply enough to run on every message, and prove with evidence that it beats your best non-fine-tuned setup. Uncertain messages will escalate to a stronger model or a human. If brand voice matters more to you, the same steps build a brand-voice classifier instead.
0:30 Analogy: a reception desk
Here is how to picture the router. It is a hospital reception desk. A trained receptionist handles most visitors quickly and correctly: that is your small fine-tuned model. When the receptionist is unsure, or the case is unusual, they call a senior nurse: that is the stronger model or a human. The receptionist does not need to know everything. They need to be fast, accurate on common cases, and honest about when to ask for help.
1:03 Steps 1–2
Step one, frame it. Write the problem statement, the list of intents with definitions, and your ladder baseline: your best prompt and model, scored on a draft test set. Fill in the fine-tuning worksheet. Step two, data. A labeling guide with priority rules, two-annotator agreement on a hundred messages, two to six thousand examples covering intents, languages and channels, PII redaction, careful synthetic top-ups, decontamination, and a split by customer or time, all documented in a data card.
1:37 Steps 3–4
Step three, train. Use LoRA or QLoRA supervised fine-tuning on a one to four billion parameter instruct model, with prompt and completion data. Track every run and choose the checkpoint by validation macro F1. Step four is optional: if one class is systematically confused and you can check it, try GRPO with a per-label grader, and keep it only if your independent test improves.
2:05 Steps 5–6
Step five, evaluate. Write your decision rule first. Compare baseline and fine-tune on the test set by slice, with bootstrap confidence intervals. Measure latency and cost per thousand messages. And run a small challenge set: ambiguous messages, messages with two intents, prompt injection attempts and out-of-scope requests. Step six, deploy and monitor: serve with multi-LoRA or a merged, quantized model, add confidence routing, a drift report, retraining triggers and rollback.
2:35 Confidence routing
Here is how the router works. The small fine-tuned model classifies each message twice, once at temperature zero and once with a little randomness, and a JSON schema forces a valid label. If both agree and the label is not other, the small model's answer stands. Otherwise the message goes to a stronger model, or to a human queue if no stronger route is allowed. If data must stay inside your boundary, make sure the stronger route is internal too.
3:10 Measure the router
Measure the router, not just the model. What share of traffic does the small model handle alone, and how accurate is it on that share? What share escalates, and what does that cost? Tune the rule until the escalated share and overall accuracy meet your targets. This is where the business case is won or lost.
3:34 Seven deliverables
Seven deliverables. The problem statement and baseline worksheet. The data card. The training record: config, versions, base model revision, data hash, curves and chosen checkpoint. The evaluation report with decision rule, slices, intervals, challenge results, routing share, latency, cost and the ship decision. The cost model with three scenarios. The deployment and monitoring plan. And a model card with intended use, limitations, languages, owner and base-model license.
4:03 Strong submission (illustrative)
What does strong look like? An illustrative example. A Lahore e-commerce enabler routes WhatsApp and email messages for forty merchants. Baseline: macro F1 of point eight four. Fine-tuned three billion model: point eight eight, with Roman Urdu at point eight six after adding seven hundred examples to thin cells. With agreement routing, the small model handles eighty-two percent of traffic alone at ninety-three percent accuracy, and cost per thousand messages falls sharply. The report also names two failure patterns and the plan to fix them.
4:40 Dry run in a day
A simple dry run before the full project. Take four intents and a hundred labeled messages. Use your best prompt with a strong model as the baseline and record its accuracy. Fine-tune a tiny model on seventy messages, test on the other thirty, and implement the two-sample agreement rule. Count how many of the thirty the small model handled alone and how many it got right. That mini version takes a day and exposes most of the practical problems you will face at full scale.
5:17 Capstone tips
A few tips for a smooth capstone. Start with the evaluation, not the training: build the test set and the baseline score in week one. Keep the first model small and the first dataset modest, so you can iterate quickly. Save every run with its config and data hash. And write the report as you go, not at the end. Most strong submissions spend more time on data and evaluation than on training itself.
5:49 FAQ: what if we don't beat the baseline?
A capstone question: what if the fine-tuned model does not beat the baseline? Then your report says so, clearly, and that is a successful capstone. You will have learned where the gap is, maybe data coverage, maybe label quality, maybe the task needs a stronger model. Write down what you would try next and what evidence would change the decision. An honest no-ship decision, backed by evaluation, is exactly the skill this course teaches.
6:21 Try this now
Try this now. Write your intent list with a one-line definition for each, plus the priority rule for the most common ambiguity. Share it with two people who handle these messages today and ask them to label the same twenty messages. Their agreement score is the first number in your capstone report.
6:44 Watch me do it (illustrative)
Watch me do it, in miniature. I take six intents and six hundred labeled messages. Baseline first: our best prompt on a strong model scores macro F1 of point eight three on a held-out set of one hundred and twenty. Then I fine-tune a small instruct model with LoRA on four hundred and eighty messages; it scores point eight six. I serve it with vLLM and run the router script on the held-out set. Two samples agree and the label is not other for ninety-eight messages; the small model is right on ninety-four of those. The other twenty-two escalate to the strong model, which gets nineteen right. Blended accuracy is higher than either alone, and the cost per thousand messages drops sharply because most never touch the strong model. I write these numbers, share handled, accuracy on that share and blended cost, straight into the evaluation report. Illustrative numbers, real workflow.
7:50 Course recap
Congratulations. You can now decide when fine-tuning is worth it, build and protect datasets, train with SFT, LoRA and QLoRA, apply DPO and GRPO, evaluate honestly, choose hosted or self-managed paths, fine-tune embeddings, model costs, and run custom models in production. Your next step is the capstone itself. Then take the final exam.
The brief
Build a support-intent router: a small, fine-tuned open-weight model that classifies incoming customer messages into your intents (and an other class), fast and cheap enough to run on every message, with a documented evaluation showing it beats your best non-fine-tuned baseline. Low-confidence messages go to a stronger model or a human.
If your organization's need is brand voice rather than routing, you may substitute a brand-voice classifier ("on-voice" / "off-voice" with reasons) using the same steps; the deliverables are identical.
Step-by-step plan
1. Frame (Module 1). Write the problem statement, the intents list with definitions, and the customization-ladder baseline: your best prompt (and model) scored on a draft test set. Complete the fine-tuning worksheet.
2. Data (Module 2).
- Labeling guide with priority rules; two-annotator agreement on 100 messages.
- 2,000–6,000 examples (real where possible), coverage matrix across intents × languages × channels.
- Redact PII; decide on synthetic top-ups for thin cells; decontaminate against test.
- Split by customer or time; data card.
3. Train (Module 3). LoRA or QLoRA SFT on a 1–4B instruct model with prompt-completion data; track runs; choose the checkpoint by validation macro-F1.
4. Optional improvement (Module 4). If one class is systematically confused and checkable, try GRPO with a per-label grader; keep it only if the independent test improves.
5. Evaluate (Module 5). Written decision rule; baseline vs fine-tune on the test set by slice with bootstrap intervals; latency and cost per 1,000 messages; a small challenge set (ambiguous, multi-intent, injection attempts).
6. Deploy and monitor (Module 6). Serve via vLLM multi-LoRA or a merged quantized model; implement confidence routing; drift report; retraining triggers; rollback.
Hands-on: confidence routing at inference
Use structured output to force a valid label, and use two cheap signals for confidence: agreement between two samples and the other label.
# router.py: small fine-tuned model first, escalate when unsure (pip install openai)
import json, os
from collections import Counter
from openai import OpenAI
small = OpenAI(base_url=os.getenv("SMALL_URL", "http://localhost:8000/v1"), api_key=os.environ["SMALL_KEY"])
large = OpenAI(base_url=os.getenv("LARGE_URL"), api_key=os.getenv("LARGE_KEY")) # stronger model or gateway
INTENTS = ["refund", "delivery", "card_blocked", "account_access", "fees", "other"]
SCHEMA = {"type": "json_schema", "json_schema": {"name": "intent", "schema": {
"type": "object", "properties": {"intent": {"type": "string", "enum": INTENTS}},
"required": ["intent"], "additionalProperties": False}}}
def classify(client, model, text, temperature):
r = client.chat.completions.create(model=model, temperature=temperature, max_tokens=20, response_format=SCHEMA,
messages=[{"role": "system", "content": "Classify the intent."},
{"role": "user", "content": text}])
return json.loads(r.choices[0].message.content)["intent"]
def route(text):
try:
votes = Counter(classify(small, "intents", text, t) for t in (0.0, 0.7))
label, n = votes.most_common(1)[0]
if n == 2 and label != "other":
return {"intent": label, "by": "small"}
except Exception as e:
print("small model error:", e)
if os.getenv("LARGE_URL"):
return {"intent": classify(large, os.environ["LARGE_MODEL"], text, 0.0), "by": "large"}
return {"intent": "other", "by": "human_queue"}
print(route("mera card ATM mein phans gaya aur ab block dikha raha hai"))Measure what share of traffic the small model handles alone and its accuracy on that share; tune the rule so the escalated share and overall accuracy meet your targets. If sensitive data must not leave your boundary, the "large" route must also be internal (see the open-weight course's hybrid routing lesson).
Deliverables
- Problem and baseline: worksheet with ladder scores.
- Data card: sources, languages, class counts, agreement, PII handling and miss rate, synthetic share.
- Training record: config file (or YAML), library versions, base model revision, data hash, curves, chosen checkpoint.
- Evaluation report (one page + appendix): decision rule; macro-F1 and per-class metrics by slice with 95% intervals; challenge-set results; latency; cost per 1,000 messages; ship/no-ship decision.
- Cost model: one-off and recurring costs, break-even under three scenarios.
- Deployment and monitoring plan: serving method, confidence routing, drift report, retraining triggers, rollback.
- Model card: intended use, limitations, languages, owner, license of base model.
Evaluation report template
| Section | Content |
|---|---|
| Decision rule | Written before training |
| Baseline | Model, prompt, score with interval |
| Fine-tune | Base, method, data version, score with interval |
| Slices | Language, channel, top/bottom classes |
| Challenge set | Ambiguous, multi-intent, injection, out-of-scope |
| Routing | Share handled by small model; accuracy on that share; escalation share |
| Ops | p50/p95 latency; cost per 1,000 messages |
| Decision | Ship / no-ship, with reasons and next steps |
Worked example: what a strong submission looks like
A Lahore-based e-commerce enabler routed WhatsApp and email messages for 40 merchants. Baseline (strong API model, tuned prompt): macro-F1 0.84. Fine-tuned 3B model with QLoRA on 4,800 examples: 0.88 overall, with Roman Urdu at 0.86 after adding 700 examples to thin cells (illustrative). With two-sample agreement routing, the small model handled 82% of traffic alone at 0.93 accuracy on that share, and escalations went to the stronger model. Cost per 1,000 messages fell sharply; break-even was about four months in the expected scenario. The report named two failure patterns (multi-intent refund-and-delivery messages; sarcasm) and the plan to address them.
Pitfalls
- Comparing against the raw base model instead of the ladder baseline.
- Letting synthetic data dominate the test set.
- No escalation path for low-confidence messages.
- Shipping without drift monitoring and rollback.
How to measure success
All seven deliverables complete, a decision rule applied honestly, and a router that meets your accuracy, escalation, latency and cost targets on real held-out data.
Key takeaways
- Frame the task and baseline first, then data, training, evaluation, deployment and monitoring
- Use prompt-completion SFT with LoRA/QLoRA on a small instruct model
- Route low-confidence messages (disagreement or "other") to a stronger model or a human
- Report slices, intervals, routing share, latency and cost in a one-page evaluation
- Ship only with monitoring, retraining triggers and rollback
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Complete the capstone: a fine-tuned support-intent (or brand-voice) model with all seven deliverables, evaluated on real held-out data against your best baseline.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.