Conversion Rate Optimization (CRO)Hypotheses and prioritisation · Lesson 9 of 20
Prioritisation with ICE, PIE and evidence scoring
Video lecture
Prioritisation with ICE, PIE and evidence scoring
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Prioritisation with ICE, PIE and evidence scoring
You have forty ideas and capacity for three tests this month. Which three? If the answer is whichever the most senior person likes, this lecture is for you. You'll learn why formal prioritisation matters, how ICE, PIE and evidence-weighted scoring work, how to deal with traffic constraints, and how to run a scoring session that the team trusts. Then you'll watch me score a backlog with a short script that makes the rules visible to everyone.
0:33 Why prioritise formally
Why prioritise formally? Because capacity is the scarcest resource in CRO. Development time, design time and above all traffic limit how many good tests you can run. Formal prioritisation makes the trade-offs visible. It reduces the influence of the loudest voice, sometimes called the HiPPO, the highest paid person's opinion. And it creates a record of why you chose what you chose, so you can improve your judgement over time.
1:03 ICE
ICE is the simplest framework. Score each idea from one to ten on impact, how much it could move the metric; confidence, how sure you are it will work; and ease, how easy it is to build and test. Multiply or average the three. It's fast and good for small teams. The weakness: impact and confidence are subjective, so scores can reflect enthusiasm more than evidence.
1:32 PIE
PIE is similar, but page-focused. Potential: how much room for improvement does this page have? Importance: how valuable is the traffic on this page? Ease: how easy is it to change? It's useful for deciding which pages to work on first. But it shares ICE's weakness. Both frameworks depend heavily on judgement, which is where evidence-weighted scoring comes in.
1:58 Evidence-weighted scoring
Evidence-weighted scoring, inspired by frameworks like PXL, replaces vague ratings with yes-or-no questions. Is the change above the fold or on a high-traffic template? Is it linked to a logged research finding? How many independent methods support it? Does it reduce anxiety or friction? Is it close to the conversion point? And how much effort is it? Each yes adds points, effort subtracts. Because the questions are factual, two people scoring the same idea tend to agree. And ideas with evidence rise to the top.
2:35 Simple example: Islamabad bakery
A simple example. A small online bakery in Islamabad has three ideas. A new homepage banner: no research link, medium effort. Adding delivery cut-off times to product pages: linked to support chats and a poll, small effort, close to purchase. And a loyalty programme: some survey support, large effort. Evidence-weighted scoring puts delivery cut-off times first, by a clear margin. The banner drops, because nobody could point to evidence it would help. The loyalty programme stays in the backlog for later research.
3:11 Which framework when?
Let's compare the three frameworks side by side, so you know when to use each. ICE works well for a small team with a short backlog, when you need to decide quickly and everyone shares context. PIE works well when you're deciding which pages or templates to focus on for the next quarter. Evidence-weighted scoring works best when the backlog is large, several people contribute ideas, and you need decisions that people trust. Many teams combine them: PIE to choose the focus pages, then evidence-weighted scoring for ideas on those pages. Whatever you choose, write the rules down and keep them stable for a few months, so scores are comparable over time.
4:00 Traffic constraints
Traffic changes everything. A brilliant idea on a page with little traffic may be untestable. Before an idea reaches the top of your list, check feasibility. How many visitors and conversions does this page get? What's the minimum worthwhile effect? How long would a test take? If the answer is six months, it's not a test candidate. Build it with monitoring, test it on a higher-traffic template, research it more, or drop it. Add a feasibility column to your backlog so this check happens every time.
4:37 Realistic example: Dubai agency session (illustrative)
Now a realistic scenario with illustrative details. A Dubai agency runs CRO for a regional electronics retailer. The backlog has forty ideas from the research sprint. In a ninety-minute session, three people score them independently with evidence-weighted questions, then discuss only the ideas where their scores differ by more than two points. Feasibility checks remove five ideas that would need months of traffic. The top three: delivery estimates on product pages, wallet payment placement in checkout, and clearer warranty information. The client's CEO had wanted a homepage video. It's in the backlog, scored transparently, with a note on what evidence would move it up.
5:22 Watch me do it: scoring script
Watch me score with a script. My backlog lives in a CSV file with one row per idea and columns for each yes-or-no question plus effort. The script holds the weights at the top, where everyone can see and challenge them. Linked research finding is worth two points. Each independent evidence source is worth one. High-traffic template, anxiety or friction reduction, and closeness to conversion are one each. Effort subtracts nothing for small, one for medium, two for large. I run it, and it prints a ranked list. The value isn't the code. It's that the rules are explicit, so arguments move from whose idea is best to whether the weights are right.
6:11 AI and prioritisation
Can AI score your backlog? It can pre-fill scores from written hypotheses and notes, which speeds up big backlogs. But impact and confidence are exactly the judgements that are easy to fake. So use AI differently: ask it to flag missing evidence and inconsistencies. For example, this hypothesis claims two evidence sources but links only one finding. Or, these two ideas are near-duplicates. Keep the final scores as a team decision.
6:42 Running a scoring session
Make scoring a team ritual. Invite the people who own research, design, development and the business goal. Score independently first, so nobody anchors on the boss. Discuss only the big disagreements. Record the final scores and the reason for any override. Rescore the top of the backlog every two weeks as new research arrives. And publish the list, so everyone can see what's next and why.
7:11 Common mistakes
Common mistakes. Scoring by enthusiasm. Letting one person override without recording why. Ignoring effort or traffic. Never rescoring as evidence changes. Treating the score as the decision, instead of an input to discussion. And letting a framework become bureaucracy. A scoring session should take ninety minutes, not a week.
7:32 Balance the portfolio
One more thing: prioritisation isn't only about the next test. Look at the whole portfolio. A healthy backlog mixes quick fixes, medium tests and a few bold bets. If every item at the top is a small copy tweak, you'll get small, hard-to-detect effects and a programme that feels busy but changes little. If every item is a bold redesign, you'll run very few tests and learn slowly. Aim for balance across pages and mechanisms too, so you're not testing the same idea on the same page again and again. Review the mix each month, and deliberately promote one bolder idea when the portfolio gets too timid.
8:19 Recap
Recap. Prioritisation makes trade-offs visible and tames opinions. ICE and PIE are fast but subjective. Evidence-weighted scoring uses factual yes-or-no questions, so ideas with evidence rise. Always check feasibility against traffic. Use AI to flag gaps, not to invent confidence. Try this now: put your backlog in a CSV with the columns from the lesson text, run the scoring script, and add a feasibility column for your top ten ideas.
Why prioritise formally
Every team has more ideas than capacity. Without a transparent method, priorities follow the loudest voice or the most senior opinion. A scoring framework makes trade-offs explicit and debatable — which is its real value. No framework is objectively "right".
ICE
Score each idea from 1 to 10 on:
- Impact — how much will this move the primary metric if it works?
- Confidence — how strong is the evidence that it will work?
- Ease — how easy is it to design, build and launch?
ICE score = the average (or product) of the three. It is fast and useful for growth teams with many small ideas, but subjective — two people can score the same idea very differently.
PIE
Designed for choosing where to test:
- Potential — how much room for improvement does this page have?
- Importance — how valuable is the traffic to this page (volume and cost)?
- Ease — how easy is it to test here (technical and political)?
PIE works well for ranking pages or templates at the start of a programme.
Evidence-weighted scoring (PXL-style)
To reduce subjectivity, some teams use binary questions — yes or no — that reward evidence and visibility. Variants of this approach are often called PXL. Example questions:
Is the change above the fold / immediately noticeable? yes = 1
Is the change noticeable within 5 seconds? yes = 2
Does it add or remove information (not just restyle)? yes = 1
Is it supported by user testing findings? yes = 1
Is it supported by qualitative feedback (surveys, VoC)? yes = 1
Is it supported by analytics data? yes = 1
Is it on a high-traffic page? yes = 1
Ease of implementation: under 4 hours = 3, 1 day = 2, 3 days = 1, longer = 0Binary questions make scoring more consistent across people.
A worked prioritisation
Illustrative backlog for an online furniture retailer, scored with ICE (1–10):
| Idea | Evidence | I | C | E | ICE (avg) |
|---|---|---|---|---|---|
| Delivery date near add-to-cart (mobile) | Survey + analytics | 7 | 8 | 8 | 7.7 |
| AR "view in your room" feature | Competitor has it | 8 | 3 | 2 | 4.3 |
| Simplify checkout from 4 to 2 steps | Funnel + recordings | 8 | 6 | 4 | 6.0 |
| New hero image on homepage | Opinion | 3 | 2 | 9 | 4.7 |
| Financing message on high-price items | Site search + sales calls | 6 | 6 | 7 | 6.3 |
The delivery-date idea ranks first: good evidence, decent impact, easy. The AR feature is exciting but low-confidence and expensive — a candidate for more research rather than immediate build. The hero image is easy but based on opinion.
Making it a team process
- Everyone can add ideas to a shared backlog — only with an evidence link.
- Score in a short weekly or fortnightly session; discuss disagreements rather than averaging silently.
- Re-score when new evidence arrives.
- Balance the roadmap: some quick wins, some bigger bets, and some research tasks.
- Record why ideas were not chosen; it prevents endless re-litigating.
Traffic constraints
Prioritisation must reflect what you can actually test. Low-traffic pages may need very large effects to reach a conclusion (see Module 5). Ideas for such pages may be better implemented directly with monitoring, or tested on a higher-traffic template.
Hands-on: an evidence-weighted scoring script
A spreadsheet works, but a small script makes the scoring rules explicit and repeatable. This version rewards evidence and penalises effort (weights are illustrative — agree your own):
import csv
WEIGHTS = {
"above_fold_or_high_traffic": 1, # change is on a high-traffic template / visible early
"addresses_research_finding": 2, # linked to a logged finding
"evidence_sources": 1, # count of independent methods (0-3), multiplied by weight
"reduces_anxiety_or_friction": 1,
"revenue_proximity": 1, # close to checkout / primary conversion
}
EFFORT_PENALTY = {"S": 0, "M": 1, "L": 2}
def score(row: dict) -> int:
s = 0
for field, w in WEIGHTS.items():
s += w * int(row.get(field, 0) or 0)
return s - EFFORT_PENALTY.get(row.get("effort", "M").strip().upper(), 1)
with open("backlog.csv", newline="", encoding="utf-8") as f:
ideas = list(csv.DictReader(f))
for idea in sorted(ideas, key=score, reverse=True):
print(f"{score(idea):>3} {idea['id']:<8} {idea['title'][:60]}")backlog.csv columns: id,title,above_fold_or_high_traffic,addresses_research_finding,evidence_sources,reduces_anxiety_or_friction,revenue_proximity,effort (yes/no fields as 1/0).
Can AI score your backlog?
AI assistants can pre-fill scores from written hypotheses and research notes, which speeds up large backlogs. But impact and confidence are exactly the judgements that are easy to fake. Use AI to flag missing evidence and inconsistencies ("HYP-031 claims two evidence sources but links only one finding"), and keep final scores as a team decision.
Add a feasibility column
Before an idea reaches the top of the list, check it is testable with your traffic (see the sample size lesson). An idea that needs 400,000 visitors per variant on a page that gets 20,000 a month is a "build" or "research more" candidate, not a test.
Common mistakes
- Scoring alone and presenting scores as objective truth.
- Inflating impact for pet ideas.
- Ignoring ease, so the roadmap stalls on complex builds.
- Never revisiting scores after new research.
- Letting a senior stakeholder's idea skip the queue without evidence (the "HiPPO" problem — highest paid person's opinion).
Running a scoring session
A 45-minute fortnightly session is usually enough. Before the meeting, each participant scores new ideas independently. In the meeting, reveal scores together, discuss only the ideas where scores differ widely, and agree final numbers. Record the rationale in one sentence. This prevents anchoring on the first person's opinion and keeps sessions short. Rotate who presents the evidence so the whole team builds research literacy.
Backlog template
| ID | Idea | Research IDs | Page/template | Hypothesis link | I | C | E | Score | Status |Key takeaways
- Frameworks make trade-offs explicit; none is objectively correct.
- ICE = Impact, Confidence, Ease; PIE = Potential, Importance, Ease.
- Binary evidence-based questions reduce scoring subjectivity.
- Require evidence links for every backlog idea and revisit scores.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Score at least eight backlog ideas with ICE and then with a binary evidence model, and compare how the rankings change.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.