Conversion Rate Optimization (CRO)Hypotheses and prioritisation · Lesson 9 of 20

Prioritisation with ICE, PIE and evidence scoring

Article · 11 min · 9 min lecture

Video lecture

Prioritisation with ICE, PIE and evidence scoring

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Prioritisation with ICE, PIE and evidence scoring

  • 40 ideas, 3 slots
  • Three frameworks compared
  • Traffic constraints
  • A scoring session people trust

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why prioritise formally

Every team has more ideas than capacity. Without a transparent method, priorities follow the loudest voice or the most senior opinion. A scoring framework makes trade-offs explicit and debatable — which is its real value. No framework is objectively "right".

ICE

Score each idea from 1 to 10 on:

  • Impact — how much will this move the primary metric if it works?
  • Confidence — how strong is the evidence that it will work?
  • Ease — how easy is it to design, build and launch?

ICE score = the average (or product) of the three. It is fast and useful for growth teams with many small ideas, but subjective — two people can score the same idea very differently.

PIE

Designed for choosing where to test:

  • Potential — how much room for improvement does this page have?
  • Importance — how valuable is the traffic to this page (volume and cost)?
  • Ease — how easy is it to test here (technical and political)?

PIE works well for ranking pages or templates at the start of a programme.

Evidence-weighted scoring (PXL-style)

To reduce subjectivity, some teams use binary questions — yes or no — that reward evidence and visibility. Variants of this approach are often called PXL. Example questions:

Is the change above the fold / immediately noticeable?            yes = 1
Is the change noticeable within 5 seconds?                        yes = 2
Does it add or remove information (not just restyle)?             yes = 1
Is it supported by user testing findings?                         yes = 1
Is it supported by qualitative feedback (surveys, VoC)?           yes = 1
Is it supported by analytics data?                                yes = 1
Is it on a high-traffic page?                                     yes = 1
Ease of implementation:     under 4 hours = 3, 1 day = 2, 3 days = 1, longer = 0

Binary questions make scoring more consistent across people.

A worked prioritisation

Illustrative backlog for an online furniture retailer, scored with ICE (1–10):

IdeaEvidenceICEICE (avg)
Delivery date near add-to-cart (mobile)Survey + analytics7887.7
AR "view in your room" featureCompetitor has it8324.3
Simplify checkout from 4 to 2 stepsFunnel + recordings8646.0
New hero image on homepageOpinion3294.7
Financing message on high-price itemsSite search + sales calls6676.3

The delivery-date idea ranks first: good evidence, decent impact, easy. The AR feature is exciting but low-confidence and expensive — a candidate for more research rather than immediate build. The hero image is easy but based on opinion.

Making it a team process

  1. Everyone can add ideas to a shared backlog — only with an evidence link.
  2. Score in a short weekly or fortnightly session; discuss disagreements rather than averaging silently.
  3. Re-score when new evidence arrives.
  4. Balance the roadmap: some quick wins, some bigger bets, and some research tasks.
  5. Record why ideas were not chosen; it prevents endless re-litigating.

Traffic constraints

Prioritisation must reflect what you can actually test. Low-traffic pages may need very large effects to reach a conclusion (see Module 5). Ideas for such pages may be better implemented directly with monitoring, or tested on a higher-traffic template.

Hands-on: an evidence-weighted scoring script

A spreadsheet works, but a small script makes the scoring rules explicit and repeatable. This version rewards evidence and penalises effort (weights are illustrative — agree your own):

import csv

WEIGHTS = {
    "above_fold_or_high_traffic": 1,   # change is on a high-traffic template / visible early
    "addresses_research_finding": 2,   # linked to a logged finding
    "evidence_sources": 1,             # count of independent methods (0-3), multiplied by weight
    "reduces_anxiety_or_friction": 1,
    "revenue_proximity": 1,            # close to checkout / primary conversion
}
EFFORT_PENALTY = {"S": 0, "M": 1, "L": 2}

def score(row: dict) -> int:
    s = 0
    for field, w in WEIGHTS.items():
        s += w * int(row.get(field, 0) or 0)
    return s - EFFORT_PENALTY.get(row.get("effort", "M").strip().upper(), 1)

with open("backlog.csv", newline="", encoding="utf-8") as f:
    ideas = list(csv.DictReader(f))

for idea in sorted(ideas, key=score, reverse=True):
    print(f"{score(idea):>3}  {idea['id']:<8} {idea['title'][:60]}")

backlog.csv columns: id,title,above_fold_or_high_traffic,addresses_research_finding,evidence_sources,reduces_anxiety_or_friction,revenue_proximity,effort (yes/no fields as 1/0).

Can AI score your backlog?

AI assistants can pre-fill scores from written hypotheses and research notes, which speeds up large backlogs. But impact and confidence are exactly the judgements that are easy to fake. Use AI to flag missing evidence and inconsistencies ("HYP-031 claims two evidence sources but links only one finding"), and keep final scores as a team decision.

Add a feasibility column

Before an idea reaches the top of the list, check it is testable with your traffic (see the sample size lesson). An idea that needs 400,000 visitors per variant on a page that gets 20,000 a month is a "build" or "research more" candidate, not a test.

Common mistakes

  • Scoring alone and presenting scores as objective truth.
  • Inflating impact for pet ideas.
  • Ignoring ease, so the roadmap stalls on complex builds.
  • Never revisiting scores after new research.
  • Letting a senior stakeholder's idea skip the queue without evidence (the "HiPPO" problem — highest paid person's opinion).

Running a scoring session

A 45-minute fortnightly session is usually enough. Before the meeting, each participant scores new ideas independently. In the meeting, reveal scores together, discuss only the ideas where scores differ widely, and agree final numbers. Record the rationale in one sentence. This prevents anchoring on the first person's opinion and keeps sessions short. Rotate who presents the evidence so the whole team builds research literacy.

Backlog template

| ID | Idea | Research IDs | Page/template | Hypothesis link | I | C | E | Score | Status |

Key takeaways

  • Frameworks make trade-offs explicit; none is objectively correct.
  • ICE = Impact, Confidence, Ease; PIE = Potential, Importance, Ease.
  • Binary evidence-based questions reduce scoring subjectivity.
  • Require evidence links for every backlog idea and revisit scores.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. What do the letters in PIE stand for?
  2. An idea scores high impact but has no evidence beyond 'a competitor does it'. How should Confidence be scored?
  3. What is the main advantage of binary evidence questions in PXL-style scoring?
  4. Your top-scoring idea is on a page with 20,000 visitors a month, but a test would need about 400,000 visitors per variant. What should you do?

Put it into practice

Score at least eight backlog ideas with ICE and then with a binary evidence model, and compare how the rankings change.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.