Gemini, Microsoft Copilot, Perplexity & the AI Tool LandscapeThe AI tool landscape · Lesson 2 of 19
How to evaluate any AI tool in 30 minutes
Video lecture
How to evaluate any AI tool in 30 minutes
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Evaluate any AI tool in 30 minutes
Every AI vendor has a stunning demo. None of those demos are your job. In this lecture you will learn how to evaluate any AI tool in thirty minutes on your own tasks, with a fair test, an honesty check that predicts real world risk, and a scorecard you can defend to your manager or client.
0:24 Why a method matters
Why use a method at all, instead of trying a tool and seeing how it feels? Because first impressions are shaped by polish, speed and confidence, not accuracy. A tool can feel brilliant for a week and quietly invent facts in the one report that matters. Think of it like test driving a car. A spin around the block is fun, but you also want to know about safety, fuel costs and what happens in the rain. The sprint is your structured test drive.
1:01 Why demos mislead
Demos are selected to shine. They use clean inputs and simple questions. The only evaluation that matters is how the tool performs on your recurring tasks, with your kind of material, under your data rules. A demo answers the question, can this tool ever do something impressive? You need the answer to a different question, will it do my work well, week after week?
1:29 Eight criteria
Use eight criteria. Task fit, does it do the job you actually have. Output quality. Grounding, does it cite sources or your files, and admit when it does not know. Integration, does it work where your work lives, through connectors or MCP. Privacy and compliance, including training use, retention and admin controls. Cost per useful output, including checking time. Reliability. And portability, can you export your prompts, files and outputs if you leave.
2:01 The sprint: minutes 0 to 20
Here is the sprint. In the first five minutes, pick three real, non confidential tasks, one writing, one analysis and one research or fact task. Then for fifteen minutes, run identical prompts with identical attachments in each tool. Paste the outputs into a document labelled A, B and C, so you can score blind without knowing which tool wrote what.
2:27 The honesty test
Minutes twenty to twenty five, the honesty test. Ask something where the correct answer is I do not know, or not in the document, such as a figure that is not in the attached report. A tool that invents an answer here will invent answers in your real work. This single test predicts risk better than any benchmark.
2:52 Score and note
In the last five minutes, score each output from one to five on accuracy, instruction following, usefulness and effort to fix, and note anything alarming, like an invented citation or an ignored constraint. Effort to fix matters most, because it is the hidden cost of every AI output.
3:13 Make it fair
Keep it fair. Compare comparable tiers, not one vendor's free plan with another's premium model. Turn on equivalent features, like search on in all tools or none. Repeat important tests, because outputs vary. And watch your own bias toward a familiar interface or a longer answer.
3:33 Data rules while testing
Until a tool is approved for a data class, test only with public or anonymised material. It is only a test is exactly how client contracts end up in unapproved tools. Before moving to real data, check the business terms, admin controls and data processing agreement.
3:53 Simple example
A simple example. Take a five page company policy and ask two tools to summarise it in five bullets. Paste the outputs as A and B and score them blind. A scores four, B scores three. Then ask both, what does the policy say about remote work from abroad, which the policy does not cover. A says it is not stated. B invents a rule. That single honesty test tells you which tool to trust with factual work, more clearly than any leaderboard.
4:29 Worked example: a Dubai research tool choice
A Dubai agency compared three research tools for competitor scans. Three tasks, blind scoring by two account managers, and an honesty test asking for a competitor's revenue, which is not publicly disclosed. One tool presented an estimate as fact, the others flagged it as unavailable. Quality scores were close, so the decision came down to integration with their existing Workspace and the data terms. They re run the sprint every six months. Six months later they reran the same sprint after a major model release. The scores shifted, and the tool that had failed the honesty test now passed it. Because they had kept the scorecard, the comparison took an hour, and they could make the switch on evidence rather than on headlines. That is the quiet advantage of a repeatable test.
5:26 Try this now
Try this now. Set a thirty minute timer. Pick three real, non confidential tasks and one honesty test question whose correct answer is not in the material. Run them in your current assistant and one alternative, paste the outputs as A and B, and score them blind on accuracy, instructions, usefulness and effort to fix. Finish by writing three sentences, which tool you recommend for which job, and why. Save the scorecard, because you will want it again in six months.
6:01 Common mistakes
Four common evaluation mistakes. Rewarding the longest, most confident answer rather than the most accurate. Testing once and generalising, when outputs vary from run to run. Forgetting that checking and fixing time is part of the cost. And using real client data in a trial because it is only a test. Avoid those, and your evaluation will be better than most vendor comparisons you read online.
6:30 Watch me do it, part 1
Let me run the sprint for the Dubai agency's research tool choice. I prepare three tasks. A competitor scan for a coffee chain, a summary of a public industry report, and a current fact about a platform policy. Plus one honesty test, what was the coffee chain's revenue last year, which is not publicly disclosed. I paste identical prompts, with the same attached report, into both tools. Then I copy the outputs into one document and label them A and B in random order, so whoever scores them does not know which tool wrote which.
7:11 Watch me do it, part 2
A colleague scores both blind. Accuracy, instructions, usefulness and minutes to fix. On the three tasks, the totals are close, eighteen versus seventeen. Then the honesty test. A says the revenue is not publicly disclosed and suggests where estimates might be found. B presents a precise figure as fact. That decides the factual work. My three sentence recommendation says, use A for research and client facing facts, B is fine for brainstorming only, and re run the sprint in six months. The scorecard goes in the shared drive with the date.
7:51 Recap and next step
Recap. Evaluate on your own tasks using eight criteria. Run identical prompts, score blind, include an honesty test, keep the comparison fair, and test only with public or anonymised data until a tool is approved. Your next step: run the thirty minute sprint comparing your main assistant with one alternative on three real tasks, and write a three sentence recommendation.
Evaluate on your tasks, not demos
Vendor demos are chosen to shine. The only evaluation that matters is how a tool performs on your recurring tasks, with your kind of material, under your data rules. A focused 30-minute sprint gets you most of the way.
Eight criteria
| Criterion | Question |
|---|---|
| Task fit | Does it do the job I actually have? |
| Output quality | Accurate, complete, well-structured, in the right tone? |
| Grounding | Does it cite sources or my files, and admit when it does not know? |
| Integration | Does it work where my work lives (suite, CRM, browser), via connectors or MCP? |
| Privacy and compliance | Training use, retention, residency, admin controls, contract terms? |
| Cost | Price per useful output, including human checking time? |
| Reliability | Consistent results, uptime, rate limits, support? |
| Portability | Can I export prompts, files and outputs if I leave? |
The 30-minute sprint
Minutes 0–5: choose three tasks. One writing task, one analysis task, one research or fact task, all real and non-confidential.
Minutes 5–20: run identical prompts in each tool with identical attachments. Paste outputs into a document labelled A, B, C so you can score blind.
Minutes 20–25: include an honesty test. Ask something whose correct answer is "I don't know" or "not in the document" (for example, a figure that is not in the attached report). Tools that invent an answer fail this test, and it predicts real-world risk.
Minutes 25–30: score each output 1–5 on accuracy, instruction-following, usefulness and effort-to-fix. Note anything alarming (invented citations, ignored constraints).
Scorecard (1-5)
Task | Tool | Accuracy | Instructions | Usefulness | Effort to fix | NotesMake the test fair
- Same prompt, same files, same time of day if usage limits apply.
- Use each tool's comparable tier (don't compare a free tier with another vendor's premium model).
- Repeat important tests; outputs vary run to run.
- Turn on equivalent features (search on in all or none).
- Watch for your own bias toward a familiar interface or longer answers.
Data rules during evaluation
Until a tool is approved for a data class, test with public or anonymised material only. "It's only a test" is how client contracts end up in unapproved tools. Check the vendor's business terms, admin controls and data processing agreement before moving to real data.
Worked example: choosing a research tool
An agency in Dubai compares Perplexity, Gemini Deep Research and ChatGPT deep research for competitor scans. Three tasks, blind scoring by two account managers, one honesty test (asking for a competitor's revenue that is not publicly disclosed). One tool presented an estimate as fact; the others flagged it as unavailable. Scores were close on quality, so the decision came down to integration (existing Workspace) and data terms. They re-run the sprint every six months.
Hands-on
Run the 30-minute sprint comparing your current primary assistant with one other tool on three real, non-confidential tasks, including an honesty test. Complete the scorecard and write a three-sentence recommendation.
A simple decision rule after scoring
- If one tool wins clearly on accuracy and effort-to-fix, and passes the honesty test, choose it, subject to data terms.
- If scores are close, decide on integration (where your work lives) and data terms (what you are allowed to send), then price.
- If a tool fails the honesty test on anything important, do not use it for factual work, however good its writing is.
- Record the decision, the date and the scorecard. Re-run the sprint every six months or after a major model release.
Scaling the sprint for a team decision
For a team purchase, ask three people from different roles to run the same three tasks independently, then compare scores. Disagreement between reviewers is useful information: it often reveals that different roles need different tools, or that the rubric needs sharpening. Add one task that uses the integration you care about (for example, drafting from a shared drive) so the test reflects real work.
Pitfalls
- Rewarding the longest or most confident answer.
- Testing once and generalising.
- Ignoring checking time in "cost".
- Using client data before approval.
How to measure success
Your tool decisions come with a scorecard, an honesty-test result and a data-terms check, and you re-test when tools change.
Key takeaways
- Evaluate on your own tasks using eight criteria: fit, quality, grounding, integration, privacy, cost, reliability and portability.
- Run identical prompts and materials across tools and score blind where possible.
- Include an 'I don't know' test and repeat important tests; outputs vary.
- Only use public or anonymised material until a tool is approved for your data.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Run the 30-minute evaluation sprint comparing your current primary assistant with one other tool on three real, non-confidential tasks. Complete the scorecard.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.