Fine-Tuning, Distillation and Custom ModelsDeploy, monitor and capstone · Lesson 16 of 16

Capstone: fine-tune a small model for support-intent routing

Article · 20 min · 8 min lecture

Video lecture

Capstone: fine-tune a small model for support-intent routing

15 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 15

Capstone: support-intent router

  • Small fine-tuned model
  • Proven against your best baseline
  • Escalation when unsure

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The brief

Build a support-intent router: a small, fine-tuned open-weight model that classifies incoming customer messages into your intents (and an other class), fast and cheap enough to run on every message, with a documented evaluation showing it beats your best non-fine-tuned baseline. Low-confidence messages go to a stronger model or a human.

If your organization's need is brand voice rather than routing, you may substitute a brand-voice classifier ("on-voice" / "off-voice" with reasons) using the same steps; the deliverables are identical.

Step-by-step plan

1. Frame (Module 1). Write the problem statement, the intents list with definitions, and the customization-ladder baseline: your best prompt (and model) scored on a draft test set. Complete the fine-tuning worksheet.

2. Data (Module 2).

  • Labeling guide with priority rules; two-annotator agreement on 100 messages.
  • 2,000–6,000 examples (real where possible), coverage matrix across intents × languages × channels.
  • Redact PII; decide on synthetic top-ups for thin cells; decontaminate against test.
  • Split by customer or time; data card.

3. Train (Module 3). LoRA or QLoRA SFT on a 1–4B instruct model with prompt-completion data; track runs; choose the checkpoint by validation macro-F1.

4. Optional improvement (Module 4). If one class is systematically confused and checkable, try GRPO with a per-label grader; keep it only if the independent test improves.

5. Evaluate (Module 5). Written decision rule; baseline vs fine-tune on the test set by slice with bootstrap intervals; latency and cost per 1,000 messages; a small challenge set (ambiguous, multi-intent, injection attempts).

6. Deploy and monitor (Module 6). Serve via vLLM multi-LoRA or a merged quantized model; implement confidence routing; drift report; retraining triggers; rollback.

Hands-on: confidence routing at inference

Use structured output to force a valid label, and use two cheap signals for confidence: agreement between two samples and the other label.

# router.py: small fine-tuned model first, escalate when unsure (pip install openai)
import json, os
from collections import Counter
from openai import OpenAI

small = OpenAI(base_url=os.getenv("SMALL_URL", "http://localhost:8000/v1"), api_key=os.environ["SMALL_KEY"])
large = OpenAI(base_url=os.getenv("LARGE_URL"), api_key=os.getenv("LARGE_KEY"))   # stronger model or gateway
INTENTS = ["refund", "delivery", "card_blocked", "account_access", "fees", "other"]
SCHEMA = {"type": "json_schema", "json_schema": {"name": "intent", "schema": {
    "type": "object", "properties": {"intent": {"type": "string", "enum": INTENTS}},
    "required": ["intent"], "additionalProperties": False}}}

def classify(client, model, text, temperature):
    r = client.chat.completions.create(model=model, temperature=temperature, max_tokens=20, response_format=SCHEMA,
                                       messages=[{"role": "system", "content": "Classify the intent."},
                                                 {"role": "user", "content": text}])
    return json.loads(r.choices[0].message.content)["intent"]

def route(text):
    try:
        votes = Counter(classify(small, "intents", text, t) for t in (0.0, 0.7))
        label, n = votes.most_common(1)[0]
        if n == 2 and label != "other":
            return {"intent": label, "by": "small"}
    except Exception as e:
        print("small model error:", e)
    if os.getenv("LARGE_URL"):
        return {"intent": classify(large, os.environ["LARGE_MODEL"], text, 0.0), "by": "large"}
    return {"intent": "other", "by": "human_queue"}

print(route("mera card ATM mein phans gaya aur ab block dikha raha hai"))

Measure what share of traffic the small model handles alone and its accuracy on that share; tune the rule so the escalated share and overall accuracy meet your targets. If sensitive data must not leave your boundary, the "large" route must also be internal (see the open-weight course's hybrid routing lesson).

Deliverables

  1. Problem and baseline: worksheet with ladder scores.
  2. Data card: sources, languages, class counts, agreement, PII handling and miss rate, synthetic share.
  3. Training record: config file (or YAML), library versions, base model revision, data hash, curves, chosen checkpoint.
  4. Evaluation report (one page + appendix): decision rule; macro-F1 and per-class metrics by slice with 95% intervals; challenge-set results; latency; cost per 1,000 messages; ship/no-ship decision.
  5. Cost model: one-off and recurring costs, break-even under three scenarios.
  6. Deployment and monitoring plan: serving method, confidence routing, drift report, retraining triggers, rollback.
  7. Model card: intended use, limitations, languages, owner, license of base model.

Evaluation report template

SectionContent
Decision ruleWritten before training
BaselineModel, prompt, score with interval
Fine-tuneBase, method, data version, score with interval
SlicesLanguage, channel, top/bottom classes
Challenge setAmbiguous, multi-intent, injection, out-of-scope
RoutingShare handled by small model; accuracy on that share; escalation share
Opsp50/p95 latency; cost per 1,000 messages
DecisionShip / no-ship, with reasons and next steps

Worked example: what a strong submission looks like

A Lahore-based e-commerce enabler routed WhatsApp and email messages for 40 merchants. Baseline (strong API model, tuned prompt): macro-F1 0.84. Fine-tuned 3B model with QLoRA on 4,800 examples: 0.88 overall, with Roman Urdu at 0.86 after adding 700 examples to thin cells (illustrative). With two-sample agreement routing, the small model handled 82% of traffic alone at 0.93 accuracy on that share, and escalations went to the stronger model. Cost per 1,000 messages fell sharply; break-even was about four months in the expected scenario. The report named two failure patterns (multi-intent refund-and-delivery messages; sarcasm) and the plan to address them.

Pitfalls

  • Comparing against the raw base model instead of the ladder baseline.
  • Letting synthetic data dominate the test set.
  • No escalation path for low-confidence messages.
  • Shipping without drift monitoring and rollback.

How to measure success

All seven deliverables complete, a decision rule applied honestly, and a router that meets your accuracy, escalation, latency and cost targets on real held-out data.

Key takeaways

  • Frame the task and baseline first, then data, training, evaluation, deployment and monitoring
  • Use prompt-completion SFT with LoRA/QLoRA on a small instruct model
  • Route low-confidence messages (disagreement or "other") to a stronger model or a human
  • Report slices, intervals, routing share, latency and cost in a one-page evaluation
  • Ship only with monitoring, retraining triggers and rollback

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. In the capstone router, which signals trigger escalation?
  2. What is the correct baseline for the capstone evaluation?
  3. Why report the share of traffic handled by the small model alone?

Put it into practice

Complete the capstone: a fine-tuned support-intent (or brand-voice) model with all seven deliverables, evaluated on real held-out data against your best baseline.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.