---
title: "Deploying custom models and monitoring for drift"
description: "Shipping is the start of the work A fine-tuned model is a snapshot of your data at one moment. Customers change how they write, products launch, policies…"
url: https://optimizeall.com/learn/fine-tuning-and-custom-models/deployment-monitoring-and-drift
updated: 2026-10-05
---

Fine-Tuning, Distillation and Custom Models · Deploy, monitor and capstone · lesson 15 of 16 · 16 min

# Deploying custom models and monitoring for drift

## Shipping is the start of the work

A fine-tuned model is a snapshot of your data at one moment. Customers change how they write, products launch, policies change, new languages appear, and the base model you built on will eventually be deprecated. Without monitoring, quality decays quietly. This lesson covers how to deploy custom models, version them, watch them, and decide when to retrain.

## Deployment options

| Option | How | Good for |
|---|---|---|
| **Multi-LoRA serving** | One base model plus many adapters loaded on demand (for example vLLM `--enable-lora --lora-modules name=path`, with limits on concurrent adapters and max rank) | Many tasks/clients on one GPU; fast adapter updates |
| **Merged + quantized** | Merge adapter into a 16-bit base, then quantize (GGUF for llama.cpp/Ollama, AWQ/FP8 for vLLM) | Single task, simplest runtime, edge devices |
| **Ollama with an adapter** | Modelfile with `FROM <base>` and `ADAPTER <path>` (supported adapter formats per Ollama docs) | Local and small-team deployments |
| **Hosted custom model** | Provider endpoint for your tuned model | Closed-model fine-tunes; no GPU ops |

Whichever you choose, the serving stack must use **the same chat template, system prompt and preprocessing** as training.

## Versioning and release

- Give every model a version: `intents-2026-09-v3` = base model revision + adapter hash + data version.
- Store adapters and merged models in a **model registry** (for example MLflow Model Registry or a private Hugging Face Hub repo) with a **model card**: base model and license, data card link, evaluation report, intended use, limitations, owner.
- Release with **champion/challenger**: the current model (champion) keeps serving while the new one (challenger) runs in shadow or on a small canary share.
- Keep **one-click rollback**: the previous version stays deployable, and routing is a config change.

## What to monitor

**1. Input drift (is traffic changing?)**

- Language and script mix, channel mix, message length distribution.
- Share of unknown or rare tokens, new product names.
- Embedding distribution shift: embed a daily sample and compare with the training distribution (for example distance between centroids, or a simple classifier that tries to tell "training" from "today"; if it succeeds easily, inputs have drifted).

**2. Output drift (is behavior changing?)**

- Label distribution (for classifiers): a sudden rise in "other" often signals a new intent.
- Format validity rate, length, refusal rate.
- Confidence proxies: agreement between two samples, or margin between top labels where available.

**3. Quality (is it still right?)**

- Weekly **human review of a random sample** (for example 100 items) against the rubric.
- **Implicit feedback:** agent re-labels, edits to drafts, customer re-contacts.
- Scheduled re-runs of the **frozen test set** plus a **fresh labeled sample** of recent traffic.

**4. Operations**

- Latency p50/p95, error rate, throughput, cost per 1,000 requests, GPU utilization.

## Hands-on: a simple drift report

```python
# drift_report.py: compare this week's inputs and predictions with the training baseline (pip install pandas numpy)
import json
import numpy as np
import pandas as pd

train = pd.read_json("train_meta.jsonl", lines=True)        # columns: lang, channel, length, label
week = pd.read_json("prod_week.jsonl", lines=True)          # same columns; label = model prediction

def psi(expected, actual, eps=1e-6):
    """Population Stability Index across categories: > 0.2 is a common 'investigate' threshold."""
    cats = sorted(set(expected) | set(actual))
    e = expected.value_counts(normalize=True).reindex(cats, fill_value=0) + eps
    a = actual.value_counts(normalize=True).reindex(cats, fill_value=0) + eps
    return float(((a - e) * np.log(a / e)).sum())

report = {
    "lang_psi": psi(train.lang, week.lang),
    "channel_psi": psi(train.channel, week.channel),
    "label_psi": psi(train.label, week.label),
    "share_other": float((week.label == "other").mean()),
    "median_length_train": float(train.length.median()),
    "median_length_week": float(week.length.median()),
}
alerts = [k for k, v in report.items() if k.endswith("_psi") and v > 0.2]
print(json.dumps(report, indent=2)); print("ALERTS:", alerts or "none")
```

Thresholds are conventions, not laws; calibrate them on a few months of your own data.

## When to retrain

Define **triggers** in advance:

- Quality on the weekly sample drops below the agreed bar.
- Drift alerts persist for two or more weeks.
- A new intent, product or language appears with meaningful volume.
- The base model is scheduled for deprecation or a better base model passes your evaluation.
- Scheduled cadence (for example quarterly) even without alarms.

Retraining reuses the pipeline: add newly labeled data (especially from failures), re-run the same evaluation plan and decision rule, release via champion/challenger.

## Worked example: a Jeddah retailer during Ramadan

The intent classifier ran well for months. In the first week of Ramadan, the share of "other" jumped and channel mix shifted to late-night WhatsApp messages about delivery slots and Eid gift wrapping. The drift report flagged label and channel PSI above threshold; the weekly review confirmed a new intent. The team labeled 700 examples, retrained, passed shadow mode in three days and released. They now schedule a pre-season data refresh before Ramadan and White Friday campaigns.

## Pitfalls

- No rollback path.
- Monitoring only latency and errors, not quality.
- Letting the base model reach end-of-life without a migration plan.
- Retraining without the same evaluation rule, so regressions slip in.

## How to measure success

Every production model has a registry entry, model card and rollback; a weekly drift and quality report exists; retraining triggers are written down and have been exercised at least once.

## Video lecture: Deploying custom models and monitoring for drift

Lecture coming soon · 15 chapters · about 8 minutes. Read the full transcript below.

1. Deploy and monitor
2. Analogy: a new employee
3. Deployment options
4. Versioning and release
5. Monitor four things
6. Detecting drift
7. Hands-on: drift_report.py
8. Retraining triggers
9. Worked example: Jeddah retailer
10. Simple example: new eSIM offer
11. Ownership
12. FAQ: how often to retrain?
13. Try this now
14. Watch me do it
15. Recap

## Lecture transcript

### Deploy and monitor

Launch day is exciting. Month four is where custom models quietly fail. Customers start writing differently, a new product launches, a festive season changes everything, and nobody notices until complaints pile up. In this lesson you will learn how to deploy custom models, version and release them safely, monitor for drift and quality, and decide when to retrain, before your users tell you.

### Analogy: a new employee

Here is an analogy. A fine-tuned model in production is like a new employee who was trained on last year's customers. At first they are excellent. But customers change, products change, and seasons change. A good manager checks in every week, reviews a sample of their work, notices new kinds of requests, and arranges refresher training when needed. Monitoring is that weekly check-in, and retraining is the refresher course.

### Deployment options

Four ways to deploy. Multi-LoRA serving puts one base model on a GPU and loads many small adapters, perfect for many tasks or clients. Merged and quantized models fold the adapter into the base and compress it, simplest for a single task or edge devices. Ollama can load a base model plus an adapter from a Modelfile for local use. And hosted endpoints serve closed-model fine-tunes. Whatever you pick, serving must use the same chat template, system prompt and preprocessing as training.

### Versioning and release

Version everything. A model version combines the base model revision, the adapter hash and the data version. Store models in a registry, with a model card: base and license, data card, evaluation report, intended use, limitations and owner. Release using champion and challenger: the current model keeps serving while the new one runs in shadow or on a small share of traffic. And always keep one-click rollback, so the previous version is a configuration change away.

### Monitor four things

What do you monitor? Four things. Input drift: language and channel mix, message length, new product names, and shifts in the distribution of embeddings. Output drift: the label distribution, format validity, length and refusal rates. Quality: a weekly human review of a random sample, implicit feedback like agent re-labels and edits, and scheduled re-runs of your test set plus a fresh labeled sample. And operations: latency, errors, cost and GPU utilization.

### Detecting drift

A neat trick for input drift: embed a daily sample of messages and compare it with your training data. You can compare the centers of the two clouds, or train a quick classifier to tell training from today. If that classifier succeeds easily, your inputs have drifted. For categories like language or label, the population stability index is a simple number to track, and values above about point two are a common signal to investigate. Calibrate thresholds on your own history.

### Hands-on: drift_report.py

The lesson's drift report is a short Python script. It loads metadata for training examples and this week's production traffic: language, channel, length and label. It computes the population stability index for language, channel and predicted label, the share of the other label, and median lengths, then prints alerts for anything above threshold. Run it weekly and put the output where the team will see it.

### Retraining triggers

When should you retrain? Decide the triggers in advance. Quality on the weekly sample falls below the bar. Drift alerts persist for two weeks. A new intent, product or language appears with real volume. The base model is scheduled for deprecation, or a better base passes your evaluation. Or simply a scheduled cadence, such as quarterly. Retraining reuses the pipeline, adds new labeled data especially from failures, and goes through the same evaluation rule and champion-challenger release.

### Worked example: Jeddah retailer

A worked example. A Jeddah retailer's intent classifier ran well for months. In the first week of Ramadan, the share of other jumped, and traffic shifted to late-night WhatsApp messages about delivery slots and Eid gift wrapping. The drift report flagged label and channel shifts, and the weekly review confirmed a new intent. They labeled seven hundred examples, retrained, passed shadow mode in three days and released. Now they refresh data before Ramadan and major sales seasons.

### Simple example: new eSIM offer

A simple example. A telecom's classifier routes messages into ten intents. One Monday, the share of other doubles. The weekly review shows many messages about a new eSIM offer that did not exist when the model was trained. They label two hundred eSIM messages, add a new intent, retrain, run the evaluation rule, shadow test for two days and release. Without the drift report, those customers would have bounced between teams for weeks.

### Ownership

Who owns all of this? Name a model owner for every production model: someone who reads the weekly report, approves retraining, signs off releases and decides on rollbacks. Pair them with an on-call engineer for operational incidents. Write both names on the model card. Monitoring that nobody reads is theater; a named owner with a weekly thirty-minute review is what actually keeps quality up.

### FAQ: how often to retrain?

A question from operations teams: how often should we retrain? There is no fixed answer, which is why triggers matter more than calendars. Some models run well for a year; others need updates every few weeks during seasonal peaks. Start with a quarterly review as a safety net, and let your drift reports and weekly quality samples tell you when to act sooner. Track how often triggers fire, and adjust your data collection so retraining is quick when it is needed.

### Try this now

Try this now. For one model you run or plan, write down the three numbers you would look at every Monday morning, for example the share of the other label, the weekly sample accuracy, and p ninety-five latency. Then write the threshold for each that would make you act. That is your monitoring plan in three lines.

### Watch me do it

Watch me do it. It is Monday morning. I run the drift report on last week's traffic. Language PSI is low. Channel PSI is point one five, fine. Label PSI is point two eight, above threshold, and the share of other has doubled. I open fifty messages labeled other and read them: thirty mention a new instalment payment option launched last Tuesday. That is a new intent. I check the weekly quality sample: accuracy on known intents is stable, so the model itself is fine; the world changed. I label two hundred instalment messages, add the intent, retrain with the same config, and run the evaluation rule: all conditions pass. I deploy the new adapter as challenger in shadow for three days, compare with the champion, then promote it. The old adapter stays in the registry for rollback. I update the model card and add instalments to the monitoring labels.

### Recap

Recap. Choose a deployment option that matches your tasks and runtime. Version base, adapter and data together, keep model cards, release champion versus challenger, and keep rollback one click away. Monitor input drift, output drift, quality and operations, and write retraining triggers in advance. Your next step: build a weekly drift report, define three triggers, and document your rollback procedure.

## Key takeaways

- Deploy via multi-LoRA serving, merged and quantized models, Ollama adapters or hosted endpoints
- Version models (base + adapter + data), keep model cards, and release champion/challenger with rollback
- Monitor input drift, output drift, sampled quality and operations
- Define retraining triggers in advance, including base-model deprecation
- Reuse the same evaluation plan and decision rule for every retrain

## Try it

Build a weekly drift report for a model you run (or plan), define three retraining triggers, and document the rollback procedure.

- [Previous: Cost modeling: training, inference and break-even](https://optimizeall.com/learn/fine-tuning-and-custom-models/cost-modeling-for-custom-models)
- [Next: Capstone: fine-tune a small model for support-intent routing](https://optimizeall.com/learn/fine-tuning-and-custom-models/capstone-support-intent-router)
- [All lessons of Fine-Tuning, Distillation and Custom Models](https://optimizeall.com/learn/fine-tuning-and-custom-models)
