Multimodal & Reasoning Models in PracticeChoosing models and benchmarking on your own data · Lesson 17 of 17

Benchmarking models on your own data

Article · 12 min · 8 min lecture

Video lecture

Benchmarking models on your own data

11 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 11

Benchmarking on your own data

  • Why public benchmarks aren't enough
  • Five steps
  • A provider-agnostic harness
  • Deciding honestly

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why public benchmarks are not enough

Public benchmarks and leaderboards are useful signals of general capability. But they have limits: they may not resemble your tasks, languages or data; some test items may have leaked into training data; scores can be sensitive to prompt formats; and small score differences may not be meaningful. A model that tops a leaderboard can underperform on your invoices, your customers' Arabic messages or your analysts' questions.

The answer is a private benchmark: a small, well-designed evaluation on your own data.

Step 1: define the task and success

Write down precisely what the model must do and what "good" looks like, including the dimensions that matter: accuracy, format compliance, tone, latency, cost. Decide which are hard requirements (valid JSON; no invented totals) and which are scored.

Step 2: build the dataset

  • Sample real inputs: 30 to 100 items is often enough to separate clearly different models; more to distinguish close ones.
  • Stratify: include easy, typical and hard cases, all relevant languages, and edge cases (poor scans, noisy audio, long inputs).
  • Label reference outputs or grading notes with domain experts.
  • Protect privacy: anonymise personal data, and make sure sending the data to each candidate provider is permitted.
  • Keep it private: do not publish it, so it stays uncontaminated.

Step 3: run fairly

  • Use the same prompt (or each model's best-practice variant, applied consistently) and the same settings.
  • Run multiple times where outputs vary, and report averages and spread.
  • Record latency and cost for every call, not just quality.
  • Blind the grading where humans judge: graders should not know which model produced which output.

Step 4: grade

Use the grading hierarchy from our evaluation lessons: code checks for objective criteria (exact fields, schema validity, numeric tolerance), rubric-based LLM judges calibrated against human ratings for subjective criteria, and human review for a sample.

{
  "model": "candidate-A",
  "task": "invoice-extraction",
  "items": 60,
  "field_accuracy": {"total": 0.97, "date": 0.95, "supplier": 0.98},
  "confident_errors": 2,
  "schema_valid": 1.0,
  "median_latency_s": 4.1,
  "cost_per_100_docs": "..."
}

(The figures above are illustrative, not real results.)

Step 5: decide with the whole picture

Build a comparison table across candidates: quality metrics, latency (median and slow tail), cost per task, and qualitative notes from reading outputs. Then weigh against non-benchmark factors: data terms, deployment options and vendor stability. Often the decision is not "the best" but "good enough at the lowest cost with acceptable terms".

Statistical caution

With 50 items, a difference of one or two correct answers is well within noise. Look for consistent differences across categories and repeated runs. If two models are close, prefer the one with better cost, latency or terms, or collect more data before deciding.

Keep the benchmark alive

  • Re-run when new models are released or current ones are updated.
  • Add items from production failures.
  • Version the dataset and results so comparisons over time are valid.

Worked example

A customer support team compares three models for drafting replies in English and Arabic. Their 80-item benchmark (40 per language) with a calibrated rubric shows two models tied in English, but one clearly better in Arabic tone and policy accuracy. The better Arabic model is slower. They route Arabic tickets to it and English tickets to the faster model. A quarterly re-run after a new release changes the English choice; the switch takes a day because prompts, dataset and harness were ready.

Hands-on: a provider-agnostic benchmark harness

Wrap each provider behind one function so the benchmark loop does not care which model it calls:

import json, os, time, statistics
import anthropic
from openai import OpenAI

anth, oai = anthropic.Anthropic(), OpenAI()

def call_claude(model, prompt):
    r = anth.messages.create(model=model, max_tokens=2048,
                             messages=[{"role": "user", "content": prompt}])
    return "".join(b.text for b in r.content if b.type == "text"), r.usage.input_tokens, r.usage.output_tokens

def call_openai(model, prompt):
    r = oai.responses.create(model=model, input=prompt)
    return r.output_text, r.usage.input_tokens, r.usage.output_tokens

CANDIDATES = {
    "claude-A": (call_claude, os.environ["CLAUDE_MODEL"]),
    "openai-B": (call_openai, os.environ["OPENAI_MODEL"]),
}

def grade(item, output):          # replace with code checks / calibrated judge
    return float(item["expected"].lower() in output.lower())

items = [json.loads(l) for l in open("bench/invoices.jsonl", encoding="utf-8")]
for name, (fn, model) in CANDIDATES.items():
    scores, lat, tin, tout = [], [], 0, 0
    for it in items:
        t0 = time.perf_counter()
        out, i, o = fn(model, it["prompt"])
        lat.append(time.perf_counter() - t0); tin += i; tout += o
        scores.append(grade(it, out))
    lat.sort()
    print(name, f"acc={statistics.mean(scores):.1%}", f"p50={lat[len(lat)//2]:.1f}s",
          f"p90={lat[int(len(lat)*0.9)]:.1f}s", f"tokens in/out={tin}/{tout}")

Add retries with backoff, run each candidate at least twice, and store raw outputs for blind human review. Convert token counts to cost with current prices kept in configuration.

Benchmarking multimodal tasks

For images, PDFs and audio, stratify by input quality (clean, average, poor) and report results per stratum. A model that wins on clean scans and loses badly on phone photos of receipts may be the wrong choice if most of your inputs are phone photos. For generation tasks (images, voice), use blind pairwise human ratings on your real briefs, with both orders shown.

Going further

Automate the harness: a script that runs any candidate model through the dataset, applies graders and outputs the comparison table. The initial investment is small compared with the cost of choosing the wrong model, and it turns every future model release into a quick, evidence-based decision rather than a debate.

Key takeaways

  • Public benchmarks are signals, not answers; they may not match your tasks, languages or data.
  • Build a private benchmark: define success, sample and stratify real inputs, label with experts, protect privacy.
  • Run fairly (same prompts, multiple runs, blind grading), record latency and cost, and grade with the code-judge-human hierarchy.
  • Decide with quality, cost, latency and terms together; beware small differences; keep the benchmark alive.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Why keep your private benchmark unpublished?
  2. Two models differ by one correct item on a 50-item benchmark. What should you conclude?
  3. Why blind human graders to which model produced each output?

Put it into practice

Build a 30-item private benchmark for one task. Run two models, record quality, latency and cost, and write a one-paragraph recommendation.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.