Multimodal & Reasoning Models in PracticeChoosing models and benchmarking on your own data · Lesson 17 of 17
Benchmarking models on your own data
Video lecture
Benchmarking models on your own data
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Benchmarking on your own data
A model tops a public leaderboard. Your team switches to it. Then Arabic replies get stiffer, invoice totals get less reliable, and nobody can say whether it is better or worse overall. Public benchmarks are useful signals, but they are not your data. In this lecture you will learn why public benchmarks are not enough, how to build a private benchmark in five steps, a provider-agnostic harness you can run today, how to benchmark multimodal tasks, and how to decide with statistics that do not fool you.
0:38 Limits of public benchmarks
Public benchmarks have limits. They may not resemble your tasks, languages or data. Some test items may have leaked into training data. Scores can be sensitive to prompt formats. And small score differences may not be meaningful. A model that leads a leaderboard can underperform on your invoices, your customers' Arabic messages or your analysts' questions. The answer is a private benchmark: a small, well-designed evaluation on your own data.
1:08 Steps 1 and 2
Step one: define the task and success. Write down precisely what the model must do, and the dimensions that matter: accuracy, format compliance, tone, latency and cost. Mark hard requirements, like valid JSON and no invented totals, separately from scored ones. Step two: build the dataset. Sample thirty to one hundred real inputs, which is often enough to separate clearly different models, and more to separate close ones. Stratify: easy, typical and hard cases, all relevant languages, and edge cases like poor scans, noisy audio and long inputs. Label reference outputs or grading notes with domain experts. Anonymise personal data, confirm you may send it to each provider, and keep the set private so it stays uncontaminated.
1:59 Steps 3 and 4
Step three: run fairly. Use the same prompt, or each model's best-practice variant applied consistently, and the same settings. Run more than once where outputs vary, and report averages and spread. Record latency and cost for every call, not just quality. And blind the grading where humans judge, so graders do not know which model produced which output. Step four: grade, using the hierarchy from the evaluation lessons: code checks for objective criteria, rubric-based judges calibrated against humans for subjective ones, and human review of a sample.
2:37 Hands-on harness
The lesson's hands-on harness wraps each provider behind one function, so the loop does not care which model it calls. The Claude wrapper and the OpenAI wrapper both return the text, input tokens and output tokens. The loop reads your items, runs each candidate, records latency and tokens, and grades each output. Then it prints accuracy, median and ninetieth-percentile latency, and total tokens for each candidate. Add retries with backoff, run each candidate at least twice, store raw outputs for blind review, and convert tokens to cost with current prices kept in configuration.
3:17 Multimodal benchmarks
Multimodal tasks need one extra habit: stratify by input quality. For images, PDFs and audio, report results for clean, average and poor inputs separately. A model that wins on clean scans and loses badly on phone photos of receipts may be the wrong choice if most of your inputs are phone photos. For generation tasks, such as images and voice, use blind pairwise human ratings on your real briefs, showing both orders to avoid position bias.
3:50 Step 5: decide
Step five: decide with the whole picture. Build a comparison table across candidates with quality metrics, median and slowest latency, cost per task, and qualitative notes from reading outputs. Then weigh non-benchmark factors: data terms, deployment options and vendor stability. The decision is often not the best, but good enough at the lowest cost with acceptable terms. And be statistically cautious. With fifty items, a difference of one or two correct answers is within noise. Look for consistent differences across categories and repeated runs, and if two models are close, prefer better cost, latency or terms, or collect more data.
4:33 Worked example: bilingual support
A worked example. A support team compares three models for drafting replies in English and Arabic, with an eighty-item benchmark, forty per language, and a calibrated rubric. Two models tie in English, but one is clearly better in Arabic tone and policy accuracy, though slower. So they route Arabic tickets to it and English tickets to the faster model. A quarterly re-run after a new release changes the English choice, and the switch takes a day, because prompts, dataset and harness were ready. Keep the benchmark alive: re-run on new releases, add items from production failures, and version the dataset and results.
5:17 Example 1: product descriptions
A simple worked example. You are choosing a model to write short product descriptions for your online shop. Instead of trusting a leaderboard, take thirty real products: ten easy, ten typical and ten awkward, like bundles or products with little information. Run two candidate models with the same prompt twice each. Grade with simple code checks, length and required facts present, and blind-rate a sample of ten descriptions per model yourself, without knowing which model wrote which. One model wins on the awkward products, where your old descriptions were weakest. That is the answer a leaderboard could never give you.
6:00 Example 2: complaint classification (illustrative)
Now a business scenario, with illustrative numbers. A bank in Pakistan wants to choose a model for classifying about one hundred thousand customer complaints a year into regulatory categories, in Urdu, English and Roman Urdu. They build a private benchmark of four hundred complaints, stratified by language and category, labelled by two compliance officers who resolve disagreements together. Three candidates are run twice with identical prompts. Results: candidate A scores ninety-one percent overall but eighty-two percent on Roman Urdu; candidate B scores eighty-nine percent overall but ninety percent on Roman Urdu, and costs about forty percent less. Because Roman Urdu is about a third of real complaints, and differences under two points overall are within noise, they choose B, record the decision with the data terms review, and schedule a re-run for the next major model release. Illustrative figures.
7:00 Recap
To recap. Public benchmarks are signals, not answers. Build a private benchmark: define success, sample and stratify real inputs, label with experts, and protect privacy. Run fairly with the same prompts, multiple runs and blind grading, record latency and cost, and grade with code, calibrated judges and humans. Decide with quality, cost, latency and terms together, and beware small differences. Try this now: build a thirty-item private benchmark for one task, run two models, and write a one-paragraph recommendation. That completes the course. Keep testing on your own data; it is your competitive advantage.
Why public benchmarks are not enough
Public benchmarks and leaderboards are useful signals of general capability. But they have limits: they may not resemble your tasks, languages or data; some test items may have leaked into training data; scores can be sensitive to prompt formats; and small score differences may not be meaningful. A model that tops a leaderboard can underperform on your invoices, your customers' Arabic messages or your analysts' questions.
The answer is a private benchmark: a small, well-designed evaluation on your own data.
Step 1: define the task and success
Write down precisely what the model must do and what "good" looks like, including the dimensions that matter: accuracy, format compliance, tone, latency, cost. Decide which are hard requirements (valid JSON; no invented totals) and which are scored.
Step 2: build the dataset
- Sample real inputs: 30 to 100 items is often enough to separate clearly different models; more to distinguish close ones.
- Stratify: include easy, typical and hard cases, all relevant languages, and edge cases (poor scans, noisy audio, long inputs).
- Label reference outputs or grading notes with domain experts.
- Protect privacy: anonymise personal data, and make sure sending the data to each candidate provider is permitted.
- Keep it private: do not publish it, so it stays uncontaminated.
Step 3: run fairly
- Use the same prompt (or each model's best-practice variant, applied consistently) and the same settings.
- Run multiple times where outputs vary, and report averages and spread.
- Record latency and cost for every call, not just quality.
- Blind the grading where humans judge: graders should not know which model produced which output.
Step 4: grade
Use the grading hierarchy from our evaluation lessons: code checks for objective criteria (exact fields, schema validity, numeric tolerance), rubric-based LLM judges calibrated against human ratings for subjective criteria, and human review for a sample.
{
"model": "candidate-A",
"task": "invoice-extraction",
"items": 60,
"field_accuracy": {"total": 0.97, "date": 0.95, "supplier": 0.98},
"confident_errors": 2,
"schema_valid": 1.0,
"median_latency_s": 4.1,
"cost_per_100_docs": "..."
}(The figures above are illustrative, not real results.)
Step 5: decide with the whole picture
Build a comparison table across candidates: quality metrics, latency (median and slow tail), cost per task, and qualitative notes from reading outputs. Then weigh against non-benchmark factors: data terms, deployment options and vendor stability. Often the decision is not "the best" but "good enough at the lowest cost with acceptable terms".
Statistical caution
With 50 items, a difference of one or two correct answers is well within noise. Look for consistent differences across categories and repeated runs. If two models are close, prefer the one with better cost, latency or terms, or collect more data before deciding.
Keep the benchmark alive
- Re-run when new models are released or current ones are updated.
- Add items from production failures.
- Version the dataset and results so comparisons over time are valid.
Worked example
A customer support team compares three models for drafting replies in English and Arabic. Their 80-item benchmark (40 per language) with a calibrated rubric shows two models tied in English, but one clearly better in Arabic tone and policy accuracy. The better Arabic model is slower. They route Arabic tickets to it and English tickets to the faster model. A quarterly re-run after a new release changes the English choice; the switch takes a day because prompts, dataset and harness were ready.
Hands-on: a provider-agnostic benchmark harness
Wrap each provider behind one function so the benchmark loop does not care which model it calls:
import json, os, time, statistics
import anthropic
from openai import OpenAI
anth, oai = anthropic.Anthropic(), OpenAI()
def call_claude(model, prompt):
r = anth.messages.create(model=model, max_tokens=2048,
messages=[{"role": "user", "content": prompt}])
return "".join(b.text for b in r.content if b.type == "text"), r.usage.input_tokens, r.usage.output_tokens
def call_openai(model, prompt):
r = oai.responses.create(model=model, input=prompt)
return r.output_text, r.usage.input_tokens, r.usage.output_tokens
CANDIDATES = {
"claude-A": (call_claude, os.environ["CLAUDE_MODEL"]),
"openai-B": (call_openai, os.environ["OPENAI_MODEL"]),
}
def grade(item, output): # replace with code checks / calibrated judge
return float(item["expected"].lower() in output.lower())
items = [json.loads(l) for l in open("bench/invoices.jsonl", encoding="utf-8")]
for name, (fn, model) in CANDIDATES.items():
scores, lat, tin, tout = [], [], 0, 0
for it in items:
t0 = time.perf_counter()
out, i, o = fn(model, it["prompt"])
lat.append(time.perf_counter() - t0); tin += i; tout += o
scores.append(grade(it, out))
lat.sort()
print(name, f"acc={statistics.mean(scores):.1%}", f"p50={lat[len(lat)//2]:.1f}s",
f"p90={lat[int(len(lat)*0.9)]:.1f}s", f"tokens in/out={tin}/{tout}")Add retries with backoff, run each candidate at least twice, and store raw outputs for blind human review. Convert token counts to cost with current prices kept in configuration.
Benchmarking multimodal tasks
For images, PDFs and audio, stratify by input quality (clean, average, poor) and report results per stratum. A model that wins on clean scans and loses badly on phone photos of receipts may be the wrong choice if most of your inputs are phone photos. For generation tasks (images, voice), use blind pairwise human ratings on your real briefs, with both orders shown.
Going further
Automate the harness: a script that runs any candidate model through the dataset, applies graders and outputs the comparison table. The initial investment is small compared with the cost of choosing the wrong model, and it turns every future model release into a quick, evidence-based decision rather than a debate.
Key takeaways
- Public benchmarks are signals, not answers; they may not match your tasks, languages or data.
- Build a private benchmark: define success, sample and stratify real inputs, label with experts, protect privacy.
- Run fairly (same prompts, multiple runs, blind grading), record latency and cost, and grade with the code-judge-human hierarchy.
- Decide with quality, cost, latency and terms together; beware small differences; keep the benchmark alive.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Build a 30-item private benchmark for one task. Run two models, record quality, latency and cost, and write a one-paragraph recommendation.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.