Gemini, Microsoft Copilot, Perplexity & the AI Tool LandscapeThe AI tool landscape · Lesson 2 of 19

How to evaluate any AI tool in 30 minutes

Article · 15 min · 8 min lecture

Video lecture

How to evaluate any AI tool in 30 minutes

16 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 16

Evaluate any AI tool in 30 minutes

  • Your tasks
  • Fair tests
  • Clear scores

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Evaluate on your tasks, not demos

Vendor demos are chosen to shine. The only evaluation that matters is how a tool performs on your recurring tasks, with your kind of material, under your data rules. A focused 30-minute sprint gets you most of the way.

Eight criteria

CriterionQuestion
Task fitDoes it do the job I actually have?
Output qualityAccurate, complete, well-structured, in the right tone?
GroundingDoes it cite sources or my files, and admit when it does not know?
IntegrationDoes it work where my work lives (suite, CRM, browser), via connectors or MCP?
Privacy and complianceTraining use, retention, residency, admin controls, contract terms?
CostPrice per useful output, including human checking time?
ReliabilityConsistent results, uptime, rate limits, support?
PortabilityCan I export prompts, files and outputs if I leave?

The 30-minute sprint

Minutes 0–5: choose three tasks. One writing task, one analysis task, one research or fact task, all real and non-confidential.

Minutes 5–20: run identical prompts in each tool with identical attachments. Paste outputs into a document labelled A, B, C so you can score blind.

Minutes 20–25: include an honesty test. Ask something whose correct answer is "I don't know" or "not in the document" (for example, a figure that is not in the attached report). Tools that invent an answer fail this test, and it predicts real-world risk.

Minutes 25–30: score each output 1–5 on accuracy, instruction-following, usefulness and effort-to-fix. Note anything alarming (invented citations, ignored constraints).

Scorecard (1-5)
Task | Tool | Accuracy | Instructions | Usefulness | Effort to fix | Notes

Make the test fair

  • Same prompt, same files, same time of day if usage limits apply.
  • Use each tool's comparable tier (don't compare a free tier with another vendor's premium model).
  • Repeat important tests; outputs vary run to run.
  • Turn on equivalent features (search on in all or none).
  • Watch for your own bias toward a familiar interface or longer answers.

Data rules during evaluation

Until a tool is approved for a data class, test with public or anonymised material only. "It's only a test" is how client contracts end up in unapproved tools. Check the vendor's business terms, admin controls and data processing agreement before moving to real data.

Worked example: choosing a research tool

An agency in Dubai compares Perplexity, Gemini Deep Research and ChatGPT deep research for competitor scans. Three tasks, blind scoring by two account managers, one honesty test (asking for a competitor's revenue that is not publicly disclosed). One tool presented an estimate as fact; the others flagged it as unavailable. Scores were close on quality, so the decision came down to integration (existing Workspace) and data terms. They re-run the sprint every six months.

Hands-on

Run the 30-minute sprint comparing your current primary assistant with one other tool on three real, non-confidential tasks, including an honesty test. Complete the scorecard and write a three-sentence recommendation.

A simple decision rule after scoring

  • If one tool wins clearly on accuracy and effort-to-fix, and passes the honesty test, choose it, subject to data terms.
  • If scores are close, decide on integration (where your work lives) and data terms (what you are allowed to send), then price.
  • If a tool fails the honesty test on anything important, do not use it for factual work, however good its writing is.
  • Record the decision, the date and the scorecard. Re-run the sprint every six months or after a major model release.

Scaling the sprint for a team decision

For a team purchase, ask three people from different roles to run the same three tasks independently, then compare scores. Disagreement between reviewers is useful information: it often reveals that different roles need different tools, or that the rubric needs sharpening. Add one task that uses the integration you care about (for example, drafting from a shared drive) so the test reflects real work.

Pitfalls

  • Rewarding the longest or most confident answer.
  • Testing once and generalising.
  • Ignoring checking time in "cost".
  • Using client data before approval.

How to measure success

Your tool decisions come with a scorecard, an honesty-test result and a data-terms check, and you re-test when tools change.

Key takeaways

  • Evaluate on your own tasks using eight criteria: fit, quality, grounding, integration, privacy, cost, reliability and portability.
  • Run identical prompts and materials across tools and score blind where possible.
  • Include an 'I don't know' test and repeat important tests; outputs vary.
  • Only use public or anonymised material until a tool is approved for your data.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. What is the most important rule for a fair side-by-side comparison?
  2. Why include a test where the correct answer is 'I don't know'?
  3. Which cost is most often overlooked when comparing AI tools?

Put it into practice

Run the 30-minute evaluation sprint comparing your current primary assistant with one other tool on three real, non-confidential tasks. Complete the scorecard.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.