Open-Weight and Local AI: Run, Choose and Deploy Your Own ModelsEdge AI, privacy and honest benchmarking · Lesson 14 of 16

Benchmarking responsibly: leaderboards, evals and honest comparisons

Article · 15 min · 8 min lecture

Video lecture

Benchmarking responsibly: leaderboards, evals and honest comparisons

14 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 14

Benchmarking responsibly

  • Why leaderboards mislead
  • Evaluations you can trust
  • Honest statistics

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why benchmarks mislead

Public benchmarks are useful for spotting candidates, and dangerous for final decisions:

  • Contamination. Test questions leak into training data, inflating scores.
  • Saturation. Top models cluster near the ceiling; differences are noise.
  • Mismatch. Benchmarks rarely match your task, language, format or difficulty mix.
  • Configuration drift. Scores depend on prompt format, sampling, reasoning effort, quantization and chat template. Vendor-reported numbers often use settings you will not run.
  • Arena-style preference votes capture what anonymous users like in open chat, which may reward length and style over correctness for your task.

Treat any headline number as a hypothesis to test.

What to measure instead

Build a task evaluation with four ingredients:

  1. A representative set. 50–300 real examples from your workload, stratified by type and difficulty, in the languages your users use. Keep a hidden holdout you never tune on.
  2. Clear scoring. Exact match or schema/field accuracy where possible; a written rubric (1–5 with anchors) for open-ended output; pairwise comparison when absolute scores are hard.
  3. Deployment-realistic settings. The same quantization, context limit, chat template, system prompt, temperature and runtime you will ship.
  4. Operational metrics. Latency (p50/p95 TTFT and total), throughput at target concurrency, memory, and cost per task.

LLM-as-judge, carefully

Using a strong model to grade outputs scales evaluation, but judges have known biases: position bias (favoring the first answer), verbosity bias (favoring longer answers), and self-preference (favoring outputs from their own family). Mitigate by:

  • Swapping order and averaging in pairwise comparisons.
  • Giving the judge a rubric and reference answer, asking for a short justification before the score.
  • Calibrating: have humans grade 50 items and check agreement with the judge before trusting it.
  • Using a judge from a different family than the candidates when possible.

Statistical honesty

With 100 examples, 82% vs 78% is often not a real difference. A quick way to see uncertainty is a bootstrap confidence interval, or simply paired comparison: on how many items did A beat B, and B beat A? Report intervals, not just point scores, and say how many examples you used.

Hands-on: paired comparison with bootstrap intervals

# compare.py: paired comparison of two models' per-item scores (0/1 or rubric 1-5)
import json, random

a = [json.loads(l)["score"] for l in open("model_a_scores.jsonl")]
b = [json.loads(l)["score"] for l in open("model_b_scores.jsonl")]
assert len(a) == len(b), "score the same items in the same order"

diffs = [x - y for x, y in zip(a, b)]
wins, losses = sum(d > 0 for d in diffs), sum(d < 0 for d in diffs)

random.seed(7)
boot = sorted(sum(random.choice(diffs) for _ in diffs) / len(diffs) for _ in range(5000))
lo, hi = boot[int(0.025 * len(boot))], boot[int(0.975 * len(boot))]
print(f"mean diff A-B = {sum(diffs)/len(diffs):+.3f}  95% CI [{lo:+.3f}, {hi:+.3f}]  A wins {wins}, B wins {losses}")
print("Real difference" if lo > 0 or hi < 0 else "Not distinguishable with this sample")

Tooling you can use

  • EleutherAI lm-evaluation-harness for standard academic benchmarks against local or served models.
  • promptfoo, Inspect (from the UK AI Security Institute) and similar frameworks for task-specific test suites, rubrics and model-graded checks.
  • Your serving engine's benchmark CLI for latency and throughput.

Whatever tool you use, version your datasets, prompts, settings and results together so comparisons are reproducible.

Worked example: a Karachi bank's model refresh

A bank's internal assistant runs a mid-size open model. A newer model tops several leaderboards. The team runs its 250-item evaluation (English, Urdu and Roman Urdu; policy Q&A, summarization, extraction) at the deployment quantization. Result (illustrative): the new model is better on English summarization, statistically indistinguishable on policy Q&A, and worse on Roman Urdu extraction. Latency is 20% higher. They keep the current model for extraction, pilot the new one for summarization, and log the decision with the evaluation report.

Writing an evaluation report

One page, every time:

SectionContents
QuestionWhat decision this evaluation informs
SetupModels, versions/hashes, quantization, runtime, prompts, sampling
DataSize, source, languages, strata, holdout policy
ResultsScores with 95% intervals per stratum; latency and cost
JudgeMethod, calibration agreement with humans
DecisionWhat you will do, and what would change your mind

Pitfalls

  • Declaring a winner on a 2–3 point gap with 50 examples.
  • Tuning prompts on the same set you report.
  • Comparing a model with reasoning on vs another with it off.
  • Letting the judge see which model wrote which answer.

How to measure success

Your model decisions come with a one-page evaluation report, confidence intervals, operational metrics and a reproducible configuration.

Key takeaways

  • Public benchmarks suffer contamination, saturation, mismatch and configuration drift
  • Evaluate on your own representative set at deployment settings, with a hidden holdout
  • LLM judges need rubrics, order swapping and calibration against humans
  • Report confidence intervals; small gaps on small samples are often noise
  • Write a one-page evaluation report for every model decision

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Model A scores 82% and Model B 78% on 100 items. What is the right conclusion?
  2. Which practice reduces LLM-judge position bias?
  3. Why must evaluation use deployment settings?

Put it into practice

Score two local models on 50 of your tasks, run the paired bootstrap script, and write a one-page evaluation report using the template.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.