Conversion Rate Optimization (CRO)Hypotheses and prioritisation · Lesson 10 of 20
Test, build or drop: choosing the right action
Video lecture
Test, build or drop: choosing the right action
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Test, build or drop
Here's a question that saves CRO teams months: does this idea actually need an A/B test? Many don't. Some are bugs you should just fix. Some are too small to measure with your traffic. Some are better tested as prototypes, or in ad platforms. And some should simply be dropped. In this lecture you'll learn a decision tree for test, build or drop, alternatives when traffic is low, how to build safely with feature flags and staged rollouts, and a feasibility check you can run in seconds.
0:38 Why the decision matters
Why does this decision matter? Because testing has costs. It takes development time, traffic and weeks of waiting. If you test a broken checkout fix, you're deliberately showing half your visitors a broken checkout for a month. That's not rigour. It's waste. On the other hand, if you ship a big redesign without testing, you might never know it quietly hurt conversion. The skill is choosing the right route for each idea, so testing capacity goes where it adds the most value.
1:14 The decision tree
Here's the decision tree. First: is it a bug, a broken experience or a legal requirement? Just do it, then monitor. Second: is it testable with your traffic, meaning you can reach the sample size for a worthwhile effect in about six weeks? If not, look for another route. Third: is the risk or uncertainty high, like a pricing change or a redesign? Then test, if feasible. Fourth: is the evidence strong and the risk low? Build it with monitoring. And if the evidence is weak and the idea isn't important, drop it or park it for more research.
1:57 The categories
Let's define the categories. Just do it: bugs, broken links, failing payment methods, accessibility failures, legal requirements. Nobody needs proof that a broken checkout is bad. Test: changes with real uncertainty and enough traffic, like a new pricing page layout or a different offer. Build: changes supported by strong evidence where testing isn't feasible or worthwhile, released with monitoring. Research more: promising ideas with thin evidence. And drop: ideas with no evidence and little potential. Dropping is a decision, not a failure.
2:33 Simple example: Lake District hotel
A simple example. A boutique hotel in the Lake District gets about three thousand website visits a month and a handful of direct bookings a week. Three ideas are on the table. The booking widget fails on some mobile browsers: just do it. Adding a clear cancellation policy near the booking button, supported by guest emails asking about it: build, and monitor bookings. And a complete homepage redesign: far too little traffic to test, so they run five quick prototype tests with past guests first. Nothing here needs a classic A/B test, and all three decisions are sound.
3:16 When a test is worth it
When is a test really worth it? Tests earn their cost when three things are true. The change carries real uncertainty, meaning smart people genuinely disagree about whether it will help. The potential impact matters to the business, because it sits on a high-value page or affects many customers. And the page has enough traffic to detect a worthwhile effect in a reasonable time. Pricing presentation, offer structure, checkout flow changes and major layout changes on key templates often meet all three. A wording tweak on a low-traffic page rarely does. Save your testing capacity for decisions where being wrong would be expensive.
4:01 Low-traffic alternatives
Low traffic isn't a dead end. It just changes the tools. Fix bugs and monitor. Build changes backed by strong qualitative evidence, using staged rollouts. Use prototype usability tests for big structural changes. Test messages and angles in ad or email platforms, where the sample is larger. Test bolder changes that could produce bigger effects, because big effects need smaller samples. And be patient. Sometimes the right answer is park it until you have more evidence.
4:34 Realistic example: Riyadh skincare (illustrative)
Now a realistic scenario with illustrative numbers. A skincare brand in Riyadh wants to test a new product page layout. The page converts at three percent, and the team wants to detect a ten percent relative improvement. At thirty thousand eligible visitors a week, that's around four weeks: a test. But their second idea is for a niche product page with five thousand visitors a week. The same maths says around twenty-one weeks. Not a test. Instead, they build the change behind a feature flag, roll it out to ten percent, then fifty, then everyone, watching errors, page speed and conversion. And they keep a small holdback group on the old version for a few weeks to estimate impact.
5:26 Watch me do it: feasibility check
Watch me run the feasibility check. I reuse the sample size function from the statistics module. I write a small helper: give it the baseline conversion rate, the minimum worthwhile effect, weekly visitors, the number of variants, and a maximum number of weeks I'm willing to run, say six. It calculates the sample per variant, converts it into weeks, and returns either test, with the expected duration, or not a test here, with the options. I run it for thirty thousand visitors a week: test, about four weeks. For five thousand: not a test, about twenty-one weeks. It takes seconds, and it ends a lot of arguments.
6:13 Building safely
Build doesn't mean forget. Use release practices from product teams. Feature flags, available in tools like GrowthBook, PostHog, LaunchDarkly, Optimizely and Statsig, let you switch a change on for a percentage of users and off instantly if something breaks. Staged rollouts, from ten to fifty to a hundred percent, catch bugs and speed regressions. Holdbacks keep a small share of users on the old experience to estimate impact. And annotate your analytics with the release date, so anyone analysing later knows what changed and when.
6:50 Decision records
Record every decision. A short decision record for each idea: the hypothesis, the route you chose, test, build, drop, research or just do it, and why. The expected effect and what you'll monitor. The date. And later, what actually happened. This turns your backlog into a learning library, even for ideas you never tested. It also helps new team members understand why things are the way they are.
7:20 Common mistakes
Common mistakes. Testing bug fixes. Testing on pages that could never reach significance. Shipping big redesigns without any research. Building without monitoring. Never dropping anything, so the backlog grows forever. And treating build as a lesser choice. For many businesses, well-researched builds with good monitoring are the majority of CRO work.
7:42 Recap
Recap. Not every idea needs an A/B test. Just do it for bugs and legal requirements. Test when uncertainty is real and traffic allows. Build strong-evidence changes with flags, staged rollouts, holdbacks and monitoring. Research more when evidence is thin, and drop what doesn't earn its place. Try this now: run the feasibility helper from the lesson text on your top five backlog ideas, assign each a route, and write a one-line decision record for each.
Not every idea deserves an A/B test
Experiments are powerful but costly: design, build, QA, weeks of runtime and analysis. Teams that test everything move slowly; teams that test nothing ship risky changes blindly. Use a decision framework for each prioritised idea.
The decision tree
Is it a bug, broken experience or legal/accessibility requirement?
-> YES: FIX NOW (monitor after release)
Is the right answer obvious and low risk (e.g. missing delivery info customers ask for)?
-> YES: JUST DO IT (implement and monitor key metrics)
Is there enough traffic to detect a realistic effect within ~2-6 weeks?
-> NO: Implement with before/after monitoring, test on a higher-traffic template,
or use qualitative validation (user tests, preference tests)
-> YES: Is the change risky or uncertain, or is learning valuable?
-> YES: A/B TEST
-> NO: JUST DO IT
Is the evidence weak?
-> RESEARCH MORE before investingCategories explained
| Action | When | Example |
|---|---|---|
| Fix | Clear bugs, errors, compliance issues | Payment form fails on one browser; missing cookie choices |
| Just do it | Obvious, low risk, supported by evidence | Adding FAQs that answer top support questions |
| Test | Uncertain outcome, meaningful risk, enough traffic | New pricing page layout; changing default plan |
| Research | Weak evidence but potentially large | New product configurator |
| Drop | Low evidence, low impact | "Make the logo bigger" |
Alternatives when traffic is low
Many small businesses and B2B sites do not have the traffic for classic A/B testing. Options:
- Qualitative validation — user tests of the new design versus the current one; five-second and preference tests.
- Bigger swings — test radically different approaches rather than small tweaks, since large effects are detectable with less traffic.
- Higher-funnel metrics — test against a more frequent event (e.g. form start rather than qualified lead), acknowledging the risk that it may not translate downstream.
- Before/after analysis with caution — compare periods while accounting for seasonality, campaigns and external factors; treat results as directional.
- Pooling traffic — test on a shared template across many pages.
Worked example: a boutique hotel in the UK
A 20-room hotel's booking page gets modest traffic, so an A/B test would take months to detect all but huge effects. Research shows guests call to ask about parking and breakfast. The team:
- Just does the obvious fix: adds a clear "What's included" block (breakfast, parking, check-in times).
- Fixes a mobile date picker bug discovered in recordings.
- Uses five user tests to validate a redesigned room comparison table before launching it.
- Monitors booking conversion and call volume before and after, noting that a local event may have affected the period.
Building a balanced roadmap
A healthy quarterly roadmap mixes:
- Fixes and quick wins that bank value immediately.
- Experiments that answer important uncertain questions.
- Research sprints that refill the pipeline with evidence.
Track all three types. A programme that only reports test wins undervalues fixes and research — and incentivises testing trivial things.
Hands-on: a feasibility check that decides test vs build
Reuse the sample-size function from the statistics module (lesson 5.2) to decide quickly whether an idea can be tested on a given page:
from sample_size import sample_size_per_variant # the function from the sample size lesson
def recommend(baseline: float, relative_mde: float, weekly_visitors: int, variants: int = 2,
max_weeks: int = 6) -> str:
n = sample_size_per_variant(baseline, relative_mde)
weeks = (n * variants) / weekly_visitors
if weeks <= max_weeks:
return f"TEST: ~{weeks:.1f} weeks ({n:,} per variant)"
return (f"NOT A TEST HERE: needs ~{weeks:.0f} weeks. Options: build + monitor, "
f"test a bolder change, use a higher-traffic template, or research more.")
print(recommend(baseline=0.03, relative_mde=0.10, weekly_visitors=30_000))
print(recommend(baseline=0.03, relative_mde=0.10, weekly_visitors=5_000))"Build" does not mean "forget"
When you build without a test, reduce risk with release practices borrowed from product teams:
- Feature flags (for example in GrowthBook, PostHog, LaunchDarkly, Optimizely or Statsig) let you turn a change on for a percentage of users and switch it off instantly if something breaks.
- Staged rollouts (10% → 50% → 100%) with guardrail monitoring catch bugs and performance regressions.
- Holdbacks: keep a small share of users (for example 5–10%) on the old experience for a few weeks to estimate impact after launch, if traffic allows.
- Annotate analytics with the release date so later analysis can account for it.
Low-traffic alternatives, summarised
| Situation | Better option than a classic A/B test |
|---|---|
| Clear bug or broken experience | Fix it (just do it), then monitor |
| Strong qualitative evidence, low traffic | Build, staged rollout, monitor guardrails |
| Big structural change, uncertain | Prototype usability tests first (see UX research methods) |
| Many small variants of ads or emails | Test in the ad or email platform, where the sample is larger |
| Still uncertain, low stakes | Drop, or park until more evidence arrives |
Common mistakes
- Testing obvious fixes for weeks while customers suffer.
- Running tests that can never reach a conclusion because traffic is too low.
- Declaring before/after changes as proven wins without considering other factors.
- Skipping research for big, expensive builds.
A step-by-step way to check if a test is feasible
- Find the weekly number of eligible visitors for the page or template.
- Find the baseline rate of the primary metric.
- Decide the smallest effect worth detecting (Module 5 explains this).
- Estimate the sample size and divide by weekly traffic to get the duration.
- If the duration exceeds roughly six to eight weeks, choose an alternative: a bolder change, a higher-frequency metric, a larger template, or implementation with monitoring.
When to drop an idea
Dropping ideas is healthy. Drop when evidence is weak and the cost is high, when the idea conflicts with brand or legal constraints, or when a similar test has already failed with no new evidence. Record the reason so the idea is not resurrected every quarter without fresh research.
Decision record template
| Idea ID | Decision (fix/do/test/research/drop) | Rationale | Owner | Date | Follow-up metric |Key takeaways
- Fix bugs and obvious issues immediately; test uncertain, risky or valuable changes.
- Low-traffic sites need qualitative validation, bigger swings or pooled templates.
- Before/after comparisons are directional, not proof.
- A balanced roadmap mixes fixes, experiments and research.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Take your top ten backlog ideas and assign each to fix, just do it, test, research or drop, with a one-line rationale.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.