---
title: "Evaluation before and after fine-tuning"
description: "No baseline, no proof A fine-tune is only \"better\" relative to something. The right comparison is not the raw base model; it is the best non-fine-tuned…"
url: https://optimizeall.com/learn/fine-tuning-and-custom-models/evaluation-before-and-after
updated: 2026-10-05
---

Fine-Tuning, Distillation and Custom Models · Evaluation, hosted platforms, embeddings and cost · lesson 11 of 16 · 16 min

# Evaluation before and after fine-tuning

## No baseline, no proof

A fine-tune is only "better" relative to something. The right comparison is not the raw base model; it is the **best non-fine-tuned setup** you reached on the customization ladder (good prompt, examples, structure, RAG, best model choice). If the fine-tune does not beat that on your target metric, at acceptable cost and risk, it should not ship.

## The four-part evaluation plan

1. **Target metric(s):** what the fine-tune is meant to improve (intent accuracy, field accuracy, rubric score for voice, win rate).
2. **Regression suite:** what must not get worse. General instruction following, other languages, safety refusals, formatting, tasks the model also serves.
3. **Operational metrics:** latency (p50/p95), throughput, cost per 1,000 requests, memory.
4. **Decision rule, written in advance:** for example "ship if intent accuracy improves by at least 3 points with a 95% interval above zero, no regression beyond 1 point on the safety suite, and p95 latency under 800 ms".

Writing the rule before training prevents moving the goalposts afterwards.

## Test sets

- **Held-out test set** from the same distribution as production, split by entity or time (see the data lesson), never used for tuning decisions.
- **Slices:** report results by language, channel, class and difficulty. An average can hide a collapse on Roman Urdu or on a rare but critical class.
- **Challenge set:** hand-written hard and adversarial cases (ambiguous intents, prompt-injection attempts, out-of-scope requests).
- **Fresh traffic sample:** a recent sample of real inputs, labeled, to check the model generalizes beyond the training period.

## Metrics by task type

| Task | Primary metrics |
|---|---|
| Classification | Accuracy, macro-F1 (treats rare classes fairly), per-class precision/recall, confusion matrix |
| Extraction | Per-field exact match, normalized match, JSON validity |
| Generation (voice, replies) | Rubric scores by trained reviewers or calibrated LLM judge; blind pairwise win rate |
| Ranking/retrieval | Recall@k, MRR, nDCG |
| Checkable reasoning | Pass rate under an independent grader |

## Hands-on: compare baseline and fine-tune through one harness

Serve the fine-tuned adapter next to the base model (for example with vLLM's LoRA support) and evaluate both through the same OpenAI-compatible client.

```bash
vllm serve Qwen/Qwen3-1.7B --enable-lora --lora-modules intents=out/intent-lora/final \
  --max-model-len 4096 --api-key "$VLLM_API_KEY"
# model "Qwen/Qwen3-1.7B" = base, model "intents" = base + adapter
```

```python
# eval_compare.py  (pip install openai scikit-learn)
import json, os
from openai import OpenAI
from sklearn.metrics import classification_report, f1_score

client = OpenAI(base_url=os.getenv("EVAL_URL", "http://localhost:8000/v1"), api_key=os.environ["VLLM_API_KEY"])
SYSTEM = "Classify the intent."          # identical to training
BASELINE_SYSTEM = open("baseline_prompt.txt").read()   # your best prompted setup from the ladder
test = [json.loads(l) for l in open("test.jsonl", encoding="utf-8")]

def predict(model, system, text):
    r = client.chat.completions.create(model=model, temperature=0, max_tokens=10,
                                       messages=[{"role": "system", "content": system}, {"role": "user", "content": text}])
    return (r.choices[0].message.content or "").strip()

gold = [ex["completion"][0]["content"] for ex in test]
texts = [ex["prompt"][-1]["content"] for ex in test]
for name, model, system in [("baseline", os.getenv("BASELINE_MODEL", "Qwen/Qwen3-1.7B"), BASELINE_SYSTEM),
                            ("fine-tuned", "intents", SYSTEM)]:
    preds = [predict(model, system, t) for t in texts]
    print(f"== {name}: macro-F1 {f1_score(gold, preds, average='macro'):.3f}")
    print(classification_report(gold, preds, zero_division=0))
```

For a fair baseline, `BASELINE_MODEL` can be a larger or API model reached through the same client pattern. Add the paired bootstrap from the benchmarking lesson to get confidence intervals, and repeat per language slice.

## Human and judge evaluation for generative tasks

- Train reviewers with the rubric and anchor examples; measure inter-rater agreement.
- Use **blind, randomized pairwise** comparisons (baseline vs fine-tune).
- If using an LLM judge, calibrate it against at least 50 human judgments and swap answer order.

## Online evaluation after offline success

1. **Shadow mode:** run the fine-tune on live traffic without showing outputs; compare with the current system.
2. **Canary / A/B:** route a small percentage of real users; watch business metrics (resolution rate, edits needed, complaints) and guard-rail metrics.
3. **Full rollout** with monitoring (Module 6).

## Worked example: a Dubai airline's baggage-claim triage

Baseline: a strong API model with a detailed prompt reached macro-F1 0.86 on 1,200 held-out claims (illustrative). The fine-tuned 3B model reached 0.89 overall, but slices showed Arabic claims at 0.81 versus 0.87 for the baseline. The decision rule required no slice to regress by more than 2 points, so they did not ship. They added 900 Arabic examples, retrained, reached 0.90 overall and 0.88 on Arabic, passed shadow mode for a week, and shipped with the API model as fallback for low-confidence cases.

## Pitfalls

- Comparing against the raw base model instead of your best prompted setup.
- Reporting only the average, hiding slice regressions.
- Tuning on the test set, even informally ("just one more try").
- Skipping safety and general-capability regressions.

## How to measure success

A written decision rule, a report with target, regression and operational metrics by slice with confidence intervals, and a clear ship/no-ship decision that follows the rule.

## Video lecture: Evaluation before and after fine-tuning

Lecture coming soon · 15 chapters · about 8 minutes. Read the full transcript below.

1. Evaluation before and after
2. Analogy: a new recipe
3. The right baseline
4. Four-part plan
5. Test sets
6. Metrics by task
7. Hands-on
8. Beyond offline
9. Worked example: Dubai airline (illustrative)
10. Simple example: travel agency (illustrative)
11. Automate it
12. FAQ: skip online testing?
13. Try this now
14. Watch me do it
15. Recap

## Lecture transcript

### Evaluation before and after

Every fine-tuning project ends with the same question: is it better? And most teams answer it badly, by comparing against the wrong thing, looking only at averages, or deciding what better means after seeing the results. In this lesson you will build an evaluation plan that gives an honest answer: the right baseline, the four parts of a plan, test sets and slices, a comparison harness, and online testing before full rollout.

### Analogy: a new recipe

Here is an analogy. Evaluating a fine-tune is like testing a new recipe in a restaurant. You do not compare it with a raw ingredient, you compare it with the dish you already serve. You taste it across the whole menu, not just one plate, so you know it does not ruin the side dishes. And you decide in advance what better means, so nobody moves the goalposts when the chef is attached to their creation.

### The right baseline

First, the baseline. Not the raw base model. Compare against the best setup you reached without fine-tuning: your best prompt, examples, structure, retrieval and model choice. That is what you would ship otherwise. If the fine-tune cannot beat it, at acceptable cost and risk, it should not ship. This single choice changes the conclusion of many projects.

### Four-part plan

A plan has four parts. Target metrics, what the fine-tune should improve. A regression suite, what must not get worse: general instructions, other languages, safety, formatting. Operational metrics: latency, throughput, cost per thousand requests. And a decision rule written in advance, for example, ship if accuracy improves by at least three points with a confidence interval above zero, no safety regression beyond one point, and p ninety-five latency under eight hundred milliseconds.

### Test sets

Test sets. A held-out test set from the production distribution, split by customer or time, never used for decisions along the way. Slices: results by language, channel, class and difficulty, because an average can hide a collapse on Roman Urdu or a rare critical class. A challenge set of hand-written hard and adversarial cases. And a fresh sample of recent traffic, to check the model generalizes beyond its training period.

### Metrics by task

Match metrics to the task. Classification: accuracy, and macro F1, which treats rare classes fairly, plus a confusion matrix. Extraction: per-field exact and normalized match, and JSON validity. Generation: rubric scores from trained reviewers or a calibrated judge, and blind pairwise win rate. Retrieval: recall at k and ranking metrics. Checkable reasoning: pass rate under an independent grader.

### Hands-on

The harness in the lesson serves your adapter beside the base model with vLLM, so one server answers as the base model or as base plus adapter by model name. The Python script runs your baseline prompt and your fine-tuned model over the same test set, through the same client, and prints macro F1 and a per-class report for each. Add the paired bootstrap from the benchmarking lesson for confidence intervals, and repeat it per language slice.

### Beyond offline

For generative tasks, train reviewers with the rubric and anchor examples, and measure their agreement. Use blind, randomized pairwise comparisons between baseline and fine-tune. If you use an AI judge, calibrate it against at least fifty human judgments and swap the answer order. Then go online. Shadow mode runs the fine-tune on live traffic without showing outputs. A canary or A/B test routes a small share of real users. Only then, full rollout with monitoring.

### Worked example: Dubai airline (illustrative)

A worked example, with illustrative numbers. A Dubai airline triages baggage claims. The baseline, a strong API model with a detailed prompt, scores macro F1 of point eight six. The fine-tuned three billion model scores point eight nine overall, but Arabic claims drop to point eight one against the baseline's point eight seven. Their rule said no slice may regress by more than two points, so they did not ship. They added nine hundred Arabic examples, retrained, reached point nine zero overall and point eight eight on Arabic, passed a week in shadow mode, and shipped with the API model as a fallback.

### Simple example: travel agency (illustrative)

A simple example. A travel agency's baseline is a strong API model with a good prompt, scoring ninety correct out of a hundred on intent classification. Their fine-tuned small model scores ninety-two. Before celebrating, they check slices: English ninety-four, Arabic eighty-eight, compared with ninety-one and eighty-nine for the baseline. Arabic dropped one point, within their two-point rule. Latency is five times lower. The rule says ship, so they ship, and keep watching the Arabic slice.

### Automate it

Make evaluation automatic. Put the test set, the baseline configuration, the scoring code and the decision rule in version control, and run them in your pipeline every time a model is trained. The pipeline produces the report and marks the model ship or no-ship according to the rule. Humans still review the report and the worst failures, but nobody has to remember how to run the evaluation, and nobody can quietly skip it.

### FAQ: skip online testing?

A question from teams under deadline pressure: can we skip the online test if offline results are great? You can shorten it, but do not skip it. Offline tests use the data you collected; live traffic brings new phrasing, new topics and new user behavior. Even a few days in shadow mode often reveals issues, like a new product name, that no offline set contained. A short shadow period is cheap insurance compared with a public failure.

### Try this now

Try this now. Write your decision rule in three lines before your next training run. Line one: the target metric and minimum improvement. Line two: the maximum regression allowed on any slice or safety test. Line three: the latency or cost limit. Share it with a colleague and ask them to hold you to it.

### Watch me do it

Watch me do it. I write the decision rule first, in a text file: ship if macro F1 improves by at least three points with the interval above zero, no language slice drops by more than two points, and p ninety-five latency is under eight hundred milliseconds. I commit it. Then I start vLLM with the base model and our adapter registered under the name intents. I run the comparison script: baseline prompt on the base model versus the adapter, same test set. Overall macro F1 rises five points. I run the bootstrap: interval comfortably above zero. Now slices: English up six, Urdu up four, Arabic down one. Within the rule. Latency: p ninety-five at four hundred milliseconds. I also run the safety suite and a small general-instruction test: no regressions. I paste all numbers into the report and check each line of the rule. All conditions met: ship to shadow mode for one week.

### Recap

Recap. Beat your best non-fine-tuned setup, not the raw base model. Plan targets, regressions, operations and a written decision rule. Report by slice with confidence intervals. Use blind comparisons for generation, then shadow and canary before rollout. Your next step: write your decision rule, evaluate baseline against fine-tune by slice, and make a clear ship or no-ship call.

## Key takeaways

- Compare against the best non-fine-tuned setup, not the raw base model
- Plan target metrics, regression suite, operational metrics and a written decision rule
- Report by slice (language, class, channel); averages hide regressions
- Use blind pairwise comparisons and calibrated judges for generation
- Follow offline wins with shadow and canary testing

## Try it

Write a decision rule for your fine-tune, then evaluate baseline vs fine-tuned on your test set by slice, with confidence intervals, and state ship or no-ship.

- [Previous: Reinforcement fine-tuning with graders: RFT and GRPO](https://optimizeall.com/learn/fine-tuning-and-custom-models/rft-and-grpo)
- [Next: Hosted fine-tuning: which platforms, which models, which trade-offs](https://optimizeall.com/learn/fine-tuning-and-custom-models/hosted-fine-tuning-apis)
- [All lessons of Fine-Tuning, Distillation and Custom Models](https://optimizeall.com/learn/fine-tuning-and-custom-models)
