Fine-Tuning, Distillation and Custom ModelsDeploy, monitor and capstone · Lesson 15 of 16

Deploying custom models and monitoring for drift

Article · 16 min · 8 min lecture

Video lecture

Deploying custom models and monitoring for drift

15 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 15

Deploy and monitor

  • Deployment options
  • Versioning and release
  • Drift, quality, retraining

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Shipping is the start of the work

A fine-tuned model is a snapshot of your data at one moment. Customers change how they write, products launch, policies change, new languages appear, and the base model you built on will eventually be deprecated. Without monitoring, quality decays quietly. This lesson covers how to deploy custom models, version them, watch them, and decide when to retrain.

Deployment options

OptionHowGood for
Multi-LoRA servingOne base model plus many adapters loaded on demand (for example vLLM --enable-lora --lora-modules name=path, with limits on concurrent adapters and max rank)Many tasks/clients on one GPU; fast adapter updates
Merged + quantizedMerge adapter into a 16-bit base, then quantize (GGUF for llama.cpp/Ollama, AWQ/FP8 for vLLM)Single task, simplest runtime, edge devices
Ollama with an adapterModelfile with FROM <base> and ADAPTER <path> (supported adapter formats per Ollama docs)Local and small-team deployments
Hosted custom modelProvider endpoint for your tuned modelClosed-model fine-tunes; no GPU ops

Whichever you choose, the serving stack must use the same chat template, system prompt and preprocessing as training.

Versioning and release

  • Give every model a version: intents-2026-09-v3 = base model revision + adapter hash + data version.
  • Store adapters and merged models in a model registry (for example MLflow Model Registry or a private Hugging Face Hub repo) with a model card: base model and license, data card link, evaluation report, intended use, limitations, owner.
  • Release with champion/challenger: the current model (champion) keeps serving while the new one (challenger) runs in shadow or on a small canary share.
  • Keep one-click rollback: the previous version stays deployable, and routing is a config change.

What to monitor

1. Input drift (is traffic changing?)

  • Language and script mix, channel mix, message length distribution.
  • Share of unknown or rare tokens, new product names.
  • Embedding distribution shift: embed a daily sample and compare with the training distribution (for example distance between centroids, or a simple classifier that tries to tell "training" from "today"; if it succeeds easily, inputs have drifted).

2. Output drift (is behavior changing?)

  • Label distribution (for classifiers): a sudden rise in "other" often signals a new intent.
  • Format validity rate, length, refusal rate.
  • Confidence proxies: agreement between two samples, or margin between top labels where available.

3. Quality (is it still right?)

  • Weekly human review of a random sample (for example 100 items) against the rubric.
  • Implicit feedback: agent re-labels, edits to drafts, customer re-contacts.
  • Scheduled re-runs of the frozen test set plus a fresh labeled sample of recent traffic.

4. Operations

  • Latency p50/p95, error rate, throughput, cost per 1,000 requests, GPU utilization.

Hands-on: a simple drift report

# drift_report.py: compare this week's inputs and predictions with the training baseline (pip install pandas numpy)
import json
import numpy as np
import pandas as pd

train = pd.read_json("train_meta.jsonl", lines=True)        # columns: lang, channel, length, label
week = pd.read_json("prod_week.jsonl", lines=True)          # same columns; label = model prediction

def psi(expected, actual, eps=1e-6):
    """Population Stability Index across categories: > 0.2 is a common 'investigate' threshold."""
    cats = sorted(set(expected) | set(actual))
    e = expected.value_counts(normalize=True).reindex(cats, fill_value=0) + eps
    a = actual.value_counts(normalize=True).reindex(cats, fill_value=0) + eps
    return float(((a - e) * np.log(a / e)).sum())

report = {
    "lang_psi": psi(train.lang, week.lang),
    "channel_psi": psi(train.channel, week.channel),
    "label_psi": psi(train.label, week.label),
    "share_other": float((week.label == "other").mean()),
    "median_length_train": float(train.length.median()),
    "median_length_week": float(week.length.median()),
}
alerts = [k for k, v in report.items() if k.endswith("_psi") and v > 0.2]
print(json.dumps(report, indent=2)); print("ALERTS:", alerts or "none")

Thresholds are conventions, not laws; calibrate them on a few months of your own data.

When to retrain

Define triggers in advance:

  • Quality on the weekly sample drops below the agreed bar.
  • Drift alerts persist for two or more weeks.
  • A new intent, product or language appears with meaningful volume.
  • The base model is scheduled for deprecation or a better base model passes your evaluation.
  • Scheduled cadence (for example quarterly) even without alarms.

Retraining reuses the pipeline: add newly labeled data (especially from failures), re-run the same evaluation plan and decision rule, release via champion/challenger.

Worked example: a Jeddah retailer during Ramadan

The intent classifier ran well for months. In the first week of Ramadan, the share of "other" jumped and channel mix shifted to late-night WhatsApp messages about delivery slots and Eid gift wrapping. The drift report flagged label and channel PSI above threshold; the weekly review confirmed a new intent. The team labeled 700 examples, retrained, passed shadow mode in three days and released. They now schedule a pre-season data refresh before Ramadan and White Friday campaigns.

Pitfalls

  • No rollback path.
  • Monitoring only latency and errors, not quality.
  • Letting the base model reach end-of-life without a migration plan.
  • Retraining without the same evaluation rule, so regressions slip in.

How to measure success

Every production model has a registry entry, model card and rollback; a weekly drift and quality report exists; retraining triggers are written down and have been exercised at least once.

Key takeaways

  • Deploy via multi-LoRA serving, merged and quantized models, Ollama adapters or hosted endpoints
  • Version models (base + adapter + data), keep model cards, and release champion/challenger with rollback
  • Monitor input drift, output drift, sampled quality and operations
  • Define retraining triggers in advance, including base-model deprecation
  • Reuse the same evaluation plan and decision rule for every retrain

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. The share of predictions labeled "other" doubles in a week. What is the most likely explanation?
  2. Why keep the previous model version deployable?
  3. Which deployment option lets one GPU serve many client-specific fine-tunes efficiently?

Put it into practice

Build a weekly drift report for a model you run (or plan), define three retraining triggers, and document the rollback procedure.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.