Gemini, Microsoft Copilot, Perplexity & the AI Tool LandscapeChoosing tools and building your stack · Lesson 17 of 19
Comparing outputs side by side
Video lecture
Comparing outputs side by side
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Comparing outputs side by side
Every few weeks a new model arrives with claims that it is the best ever. How do you know whether it is better for your work? In this lecture you will learn to build a reusable test set from your real tasks, score outputs with a clear rubric, use AI to help judge without being fooled, and read results so you make good decisions, not just exciting ones.
0:30 Why a test set
Why build a test set? Because impressions are unreliable. The tool you used most recently, or the one with the prettiest interface, tends to feel better. And models change often, so a comparison from last spring may be wrong today. A fixed set of real tasks with success criteria lets you compare fairly, and re run the comparison in under an hour whenever something changes. AI engineers call this an evaluation set. You can build one without writing any code.
1:05 The analogy
Think of it like a driving test. Every candidate drives the same route, under the same rules, and passes or fails on clear criteria, not on how charming they are. Your test set is the route. Your rubric is the examiner's checklist. And every AI tool takes exactly the same test.
1:27 Building the test set
Build eight to fifteen tasks from your real work. A writing task, like a client update from bullet notes. A summary of a ten page report. An analysis of a campaign CSV. A research question needing a primary source. An instruction following task with a strict format, word limit and banned words. An honesty test, a question the material cannot answer. A language task if you work in Arabic, Urdu or another language, reviewed by a native speaker. And one task unique to your job.
2:04 Criteria first
Here is the key idea. Write the success criteria before you run anything. All key numbers present and correct. Every constraint met. Says not in the material when appropriate. If you write criteria after seeing the outputs, an impressive looking answer will quietly move the goalposts, and you will reward style over substance.
2:27 The rubric
Score each output from one to five on four things. Accuracy, meaning facts, numbers and citations. Instruction following, meaning every constraint met. Usefulness, meaning could you use it with light edits. And honesty, meaning it admits uncertainty and invents nothing. Add effort to fix in minutes, which often decides real value. And score blind where you can. Paste outputs as A, B and C, and ask a colleague to score without knowing which tool wrote which.
3:00 Simple example
A simple example. Two tools summarise the same ten page report in five bullets. B's summary is longer and reads beautifully. A's is plainer. You compare both against your reference facts from the report. A has every number right. B has one revenue figure wrong. A wins, clearly, despite being less impressive at first glance. That is exactly the kind of difference a rubric catches and a quick impression misses.
3:30 AI as judge, carefully
You can use an assistant to help score large test sets, but know its biases. AI judges tend to prefer longer answers, their own style or vendor, and confident tone. So give the judge your rubric and the reference facts, ask it to quote the evidence for each score, randomise the order and hide the tool names, and spot check a sample yourself. The AI helps with the volume. You make the final accuracy calls.
4:03 Reading results
Now read the results wisely. Look at failure modes, not just averages. A tool that averages four point two but occasionally invents financial figures may be unacceptable for finance work. Ignore tiny differences, like four point one versus four point zero, which are within normal variation. Weigh effort to fix and integration alongside the scores. And re test after major model releases.
4:30 Business example: Lahore agency, 20 people
A realistic business example. A twenty person agency in Lahore builds a twelve task test set from real client work, including an Urdu task and an honesty test, and runs it on three assistants. Two account managers score blind, an AI judge scores with quoted evidence, and a sample of its scores is checked by people. Two tools are close on average, but one invents a statistic in the research task. The agency standardises on the other, documents why, and schedules a re test in six months. Six months later, the agency reran the same twelve tasks when both vendors released new models. The tool they had chosen still led on the research task, but the other had fixed its invented statistics problem. Because the test set and past results were saved, they could see the change clearly and decided to keep their choice while adding the other tool for a specific drafting task.
5:37 Common mistakes
Four common mistakes. Test sets made only of easy tasks, which every tool passes. Criteria written after seeing the outputs. Trusting an AI judge without spot checks. And deciding on averages while ignoring a dangerous failure. Avoid those, and your comparison will be better than most published reviews.
5:58 Keep the test set alive
Treat your test set like a living document. Store the tasks, inputs, success criteria and reference facts in a shared folder with a version number. When your work changes, add tasks. When one becomes irrelevant, retire it. And keep past results, so when a vendor announces a new model, you can see whether it actually improved your scores, rather than relying on memory or marketing.
6:26 Watch me do it, part 1
Let me build the Lahore agency's test set. In a sheet I list ten tasks from real client work, with columns for the input, the success criteria and the reference facts. Task six is an honesty test, a question the attached report cannot answer. Task nine is an Urdu message, to be reviewed by a native speaker. I fill every success criteria cell before running anything. Then I run all ten in two assistants, with identical prompts and files, and paste the outputs into a document as A and B, with tool names hidden.
7:07 Watch me do it, part 2
I paste the rubric, the reference facts and the outputs into the judge prompt, asking for quoted evidence and telling it not to reward length. The judge scores each row with quotes. I spot check three rows myself. In one, the judge rewarded B's longer answer despite a missed constraint, so I lower that score. Then I look for failure modes. In task four, B invented a market statistic. That single failure matters more than the averages. My summary recommends A for research and client facts, and says both are fine for internal drafting.
7:48 Recap and try this now
Recap. Build a reusable test set of eight to fifteen real tasks, write criteria first, score with the four part rubric plus effort to fix, use AI judges carefully, and decide on failure modes as well as averages. Try this now. Write ten tasks from your real work with success criteria, run them on two tools, and write a one page summary with your recommendation and the most important failure modes you found.
Why a reusable test set beats ad-hoc impressions
The 30-minute sprint (Lesson 1.2) is great for a first look. For ongoing decisions (which assistant to standardise on, whether a new model is better, whether a prompt change helped) you need a reusable test set: a fixed collection of real tasks with success criteria that you can re-run whenever tools change. It is the non-technical version of what AI engineers call an evaluation set.
Building a test set (8 to 15 tasks)
Include a mix that reflects your real work:
| Task type | Example | Success criteria |
|---|---|---|
| Writing | Client update email from bullet notes | Accurate facts, correct tone, under 150 words |
| Summarising | 10-page report → 5 bullets | All key numbers present and correct |
| Analysis | CSV of campaign results → top 3 insights | Correct calculations; caveats stated |
| Research | Current rule or fact with sources | Primary source cited; date given |
| Instruction-following | Strict format with word limit and banned words | Every constraint met |
| Honesty | Question not answerable from the material | Says "not in the material" |
| Language | Arabic or Urdu version of a short message | Native reviewer approves |
| Your specialty | A task unique to your job | Defined by you |
Write the success criteria before running anything, so impressive-looking answers cannot move the goalposts.
A scoring rubric
Score each output 1 to 5 on:
- Accuracy: facts, numbers and citations correct.
- Instruction-following: constraints met (length, format, tone, exclusions).
- Usefulness: could you use it with light edits?
- Honesty: admits uncertainty; no invented facts or sources.
Add effort-to-fix in minutes, which often decides real-world value.
Score blind where possible: paste outputs as A, B, C and ask a colleague to score without knowing the tool.
Using AI to help judge (carefully)
You can ask an assistant to score outputs against your rubric, which saves time on large test sets. Known biases of "AI as judge": preference for longer answers, for its own style or its own vendor's outputs, and for confident tone. Mitigations:
- Give the judge the rubric and the reference facts, and ask for evidence (quote the part that meets or fails each criterion).
- Randomise order; hide tool names.
- Spot-check a sample of the judge's scores yourself; you make the final accuracy calls.
You are scoring outputs against a rubric. For each output (A, B, C):
- Accuracy (1-5): compare with the reference facts below; quote any error.
- Instructions (1-5): list each constraint and whether it was met.
- Honesty (1-5): did it admit what it could not know?
Return a table and a one-line justification per score. Do not reward length.
<reference_facts>...</reference_facts>
<outputs>...</outputs>Reading the results
- Look at failure modes, not averages. A tool that averages 4.2 but occasionally invents financial figures may be unacceptable for finance work.
- Ignore tiny differences (4.1 vs 4.0); they are within run-to-run variation.
- Weigh effort-to-fix and integration alongside scores.
- Re-test after major model releases or when a vendor changes plans.
Worked example: an agency standardises its assistant
A 20-person agency in Lahore builds a 12-task test set from real client work (including an Urdu task and an honesty test) and runs it on three assistants. Two account managers score blind; an AI judge scores too, with evidence, and a sample is checked by humans. Two tools are close on average; one invents a statistic in the research task. The agency standardises on the other, documents why, and schedules a re-test in six months.
Hands-on
Build a 10-task test set from your real work with success criteria. Run it on two tools and write a one-page summary with a recommendation and the key failure modes you found.
Keeping the test set alive
Store the test set (tasks, inputs, success criteria, reference facts) in a shared folder with a version number. When your work changes, add tasks; when a task becomes irrelevant, retire it. Keep past results so you can see whether a new model or prompt actually improved things, rather than relying on memory.
Pitfalls
- Test sets made of easy tasks only.
- Criteria written after seeing the outputs.
- Trusting an AI judge without spot-checks.
- Deciding on averages while ignoring dangerous failures.
How to measure success
Your tool decisions rest on a documented test set, you can re-run it in under an hour, and your team trusts the recommendation because the method is visible.
Key takeaways
- Build a reusable test set of 8 to 15 real tasks with success criteria, honesty and instruction-following tests.
- Score with a rubric (accuracy, instruction following, usefulness, honesty), blind where possible.
- AI can help evaluate, but has biases; make final accuracy judgements yourself.
- Focus on failure modes and effort-to-fix, ignore small differences, and re-test after updates.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Build a 10-task test set from your real work with success criteria. Run it on two tools and write a one-page results summary with a recommendation.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.