Gemini, Microsoft Copilot, Perplexity & the AI Tool LandscapeChoosing tools and building your stack · Lesson 17 of 19

Comparing outputs side by side

Article · 16 min · 8 min lecture

Video lecture

Comparing outputs side by side

15 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 15

Comparing outputs side by side

  • A reusable test set
  • A clear rubric
  • Honest judging

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why a reusable test set beats ad-hoc impressions

The 30-minute sprint (Lesson 1.2) is great for a first look. For ongoing decisions (which assistant to standardise on, whether a new model is better, whether a prompt change helped) you need a reusable test set: a fixed collection of real tasks with success criteria that you can re-run whenever tools change. It is the non-technical version of what AI engineers call an evaluation set.

Building a test set (8 to 15 tasks)

Include a mix that reflects your real work:

Task typeExampleSuccess criteria
WritingClient update email from bullet notesAccurate facts, correct tone, under 150 words
Summarising10-page report → 5 bulletsAll key numbers present and correct
AnalysisCSV of campaign results → top 3 insightsCorrect calculations; caveats stated
ResearchCurrent rule or fact with sourcesPrimary source cited; date given
Instruction-followingStrict format with word limit and banned wordsEvery constraint met
HonestyQuestion not answerable from the materialSays "not in the material"
LanguageArabic or Urdu version of a short messageNative reviewer approves
Your specialtyA task unique to your jobDefined by you

Write the success criteria before running anything, so impressive-looking answers cannot move the goalposts.

A scoring rubric

Score each output 1 to 5 on:

  1. Accuracy: facts, numbers and citations correct.
  2. Instruction-following: constraints met (length, format, tone, exclusions).
  3. Usefulness: could you use it with light edits?
  4. Honesty: admits uncertainty; no invented facts or sources.

Add effort-to-fix in minutes, which often decides real-world value.

Score blind where possible: paste outputs as A, B, C and ask a colleague to score without knowing the tool.

Using AI to help judge (carefully)

You can ask an assistant to score outputs against your rubric, which saves time on large test sets. Known biases of "AI as judge": preference for longer answers, for its own style or its own vendor's outputs, and for confident tone. Mitigations:

  • Give the judge the rubric and the reference facts, and ask for evidence (quote the part that meets or fails each criterion).
  • Randomise order; hide tool names.
  • Spot-check a sample of the judge's scores yourself; you make the final accuracy calls.
You are scoring outputs against a rubric. For each output (A, B, C):
- Accuracy (1-5): compare with the reference facts below; quote any error.
- Instructions (1-5): list each constraint and whether it was met.
- Honesty (1-5): did it admit what it could not know?
Return a table and a one-line justification per score. Do not reward length.
<reference_facts>...</reference_facts>
<outputs>...</outputs>

Reading the results

  • Look at failure modes, not averages. A tool that averages 4.2 but occasionally invents financial figures may be unacceptable for finance work.
  • Ignore tiny differences (4.1 vs 4.0); they are within run-to-run variation.
  • Weigh effort-to-fix and integration alongside scores.
  • Re-test after major model releases or when a vendor changes plans.

Worked example: an agency standardises its assistant

A 20-person agency in Lahore builds a 12-task test set from real client work (including an Urdu task and an honesty test) and runs it on three assistants. Two account managers score blind; an AI judge scores too, with evidence, and a sample is checked by humans. Two tools are close on average; one invents a statistic in the research task. The agency standardises on the other, documents why, and schedules a re-test in six months.

Hands-on

Build a 10-task test set from your real work with success criteria. Run it on two tools and write a one-page summary with a recommendation and the key failure modes you found.

Keeping the test set alive

Store the test set (tasks, inputs, success criteria, reference facts) in a shared folder with a version number. When your work changes, add tasks; when a task becomes irrelevant, retire it. Keep past results so you can see whether a new model or prompt actually improved things, rather than relying on memory.

Pitfalls

  • Test sets made of easy tasks only.
  • Criteria written after seeing the outputs.
  • Trusting an AI judge without spot-checks.
  • Deciding on averages while ignoring dangerous failures.

How to measure success

Your tool decisions rest on a documented test set, you can re-run it in under an hour, and your team trusts the recommendation because the method is visible.

Key takeaways

  • Build a reusable test set of 8 to 15 real tasks with success criteria, honesty and instruction-following tests.
  • Score with a rubric (accuracy, instruction following, usefulness, honesty), blind where possible.
  • AI can help evaluate, but has biases; make final accuracy judgements yourself.
  • Focus on failure modes and effort-to-fix, ignore small differences, and re-test after updates.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Why write success criteria before running comparison tasks?
  2. What is a known bias of using an AI to judge other AI outputs?
  3. Tool X scores slightly higher on average but occasionally invents financial figures. For finance reports, what should you consider?

Put it into practice

Build a 10-task test set from your real work with success criteria. Run it on two tools and write a one-page results summary with a recommendation.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.