Conversion Rate Optimization (CRO)Hypotheses and prioritisation · Lesson 10 of 20

Test, build or drop: choosing the right action

Article · 10 min · 8 min lecture

Video lecture

Test, build or drop: choosing the right action

13 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 13

Test, build or drop

  • Not every idea needs an A/B test
  • A decision tree
  • Low-traffic alternatives
  • Building safely with flags and rollouts

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Not every idea deserves an A/B test

Experiments are powerful but costly: design, build, QA, weeks of runtime and analysis. Teams that test everything move slowly; teams that test nothing ship risky changes blindly. Use a decision framework for each prioritised idea.

The decision tree

Is it a bug, broken experience or legal/accessibility requirement?
  -> YES: FIX NOW (monitor after release)
Is the right answer obvious and low risk (e.g. missing delivery info customers ask for)?
  -> YES: JUST DO IT (implement and monitor key metrics)
Is there enough traffic to detect a realistic effect within ~2-6 weeks?
  -> NO:  Implement with before/after monitoring, test on a higher-traffic template,
          or use qualitative validation (user tests, preference tests)
  -> YES: Is the change risky or uncertain, or is learning valuable?
          -> YES: A/B TEST
          -> NO:  JUST DO IT
Is the evidence weak?
  -> RESEARCH MORE before investing

Categories explained

ActionWhenExample
FixClear bugs, errors, compliance issuesPayment form fails on one browser; missing cookie choices
Just do itObvious, low risk, supported by evidenceAdding FAQs that answer top support questions
TestUncertain outcome, meaningful risk, enough trafficNew pricing page layout; changing default plan
ResearchWeak evidence but potentially largeNew product configurator
DropLow evidence, low impact"Make the logo bigger"

Alternatives when traffic is low

Many small businesses and B2B sites do not have the traffic for classic A/B testing. Options:

  • Qualitative validation — user tests of the new design versus the current one; five-second and preference tests.
  • Bigger swings — test radically different approaches rather than small tweaks, since large effects are detectable with less traffic.
  • Higher-funnel metrics — test against a more frequent event (e.g. form start rather than qualified lead), acknowledging the risk that it may not translate downstream.
  • Before/after analysis with caution — compare periods while accounting for seasonality, campaigns and external factors; treat results as directional.
  • Pooling traffic — test on a shared template across many pages.

Worked example: a boutique hotel in the UK

A 20-room hotel's booking page gets modest traffic, so an A/B test would take months to detect all but huge effects. Research shows guests call to ask about parking and breakfast. The team:

  1. Just does the obvious fix: adds a clear "What's included" block (breakfast, parking, check-in times).
  2. Fixes a mobile date picker bug discovered in recordings.
  3. Uses five user tests to validate a redesigned room comparison table before launching it.
  4. Monitors booking conversion and call volume before and after, noting that a local event may have affected the period.

Building a balanced roadmap

A healthy quarterly roadmap mixes:

  • Fixes and quick wins that bank value immediately.
  • Experiments that answer important uncertain questions.
  • Research sprints that refill the pipeline with evidence.

Track all three types. A programme that only reports test wins undervalues fixes and research — and incentivises testing trivial things.

Hands-on: a feasibility check that decides test vs build

Reuse the sample-size function from the statistics module (lesson 5.2) to decide quickly whether an idea can be tested on a given page:

from sample_size import sample_size_per_variant   # the function from the sample size lesson

def recommend(baseline: float, relative_mde: float, weekly_visitors: int, variants: int = 2,
              max_weeks: int = 6) -> str:
    n = sample_size_per_variant(baseline, relative_mde)
    weeks = (n * variants) / weekly_visitors
    if weeks <= max_weeks:
        return f"TEST: ~{weeks:.1f} weeks ({n:,} per variant)"
    return (f"NOT A TEST HERE: needs ~{weeks:.0f} weeks. Options: build + monitor, "
            f"test a bolder change, use a higher-traffic template, or research more.")

print(recommend(baseline=0.03, relative_mde=0.10, weekly_visitors=30_000))
print(recommend(baseline=0.03, relative_mde=0.10, weekly_visitors=5_000))

"Build" does not mean "forget"

When you build without a test, reduce risk with release practices borrowed from product teams:

  • Feature flags (for example in GrowthBook, PostHog, LaunchDarkly, Optimizely or Statsig) let you turn a change on for a percentage of users and switch it off instantly if something breaks.
  • Staged rollouts (10% → 50% → 100%) with guardrail monitoring catch bugs and performance regressions.
  • Holdbacks: keep a small share of users (for example 5–10%) on the old experience for a few weeks to estimate impact after launch, if traffic allows.
  • Annotate analytics with the release date so later analysis can account for it.

Low-traffic alternatives, summarised

SituationBetter option than a classic A/B test
Clear bug or broken experienceFix it (just do it), then monitor
Strong qualitative evidence, low trafficBuild, staged rollout, monitor guardrails
Big structural change, uncertainPrototype usability tests first (see UX research methods)
Many small variants of ads or emailsTest in the ad or email platform, where the sample is larger
Still uncertain, low stakesDrop, or park until more evidence arrives

Common mistakes

  • Testing obvious fixes for weeks while customers suffer.
  • Running tests that can never reach a conclusion because traffic is too low.
  • Declaring before/after changes as proven wins without considering other factors.
  • Skipping research for big, expensive builds.

A step-by-step way to check if a test is feasible

  1. Find the weekly number of eligible visitors for the page or template.
  2. Find the baseline rate of the primary metric.
  3. Decide the smallest effect worth detecting (Module 5 explains this).
  4. Estimate the sample size and divide by weekly traffic to get the duration.
  5. If the duration exceeds roughly six to eight weeks, choose an alternative: a bolder change, a higher-frequency metric, a larger template, or implementation with monitoring.

When to drop an idea

Dropping ideas is healthy. Drop when evidence is weak and the cost is high, when the idea conflicts with brand or legal constraints, or when a similar test has already failed with no new evidence. Record the reason so the idea is not resurrected every quarter without fresh research.

Decision record template

| Idea ID | Decision (fix/do/test/research/drop) | Rationale | Owner | Date | Follow-up metric |

Key takeaways

  • Fix bugs and obvious issues immediately; test uncertain, risky or valuable changes.
  • Low-traffic sites need qualitative validation, bigger swings or pooled templates.
  • Before/after comparisons are directional, not proof.
  • A balanced roadmap mixes fixes, experiments and research.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A form fails on one browser, blocking submissions. What is the right action?
  2. A B2B site with low traffic wants to validate a new pricing layout. Which approach is most practical?
  3. Why are before/after comparisons weaker evidence than A/B tests?
  4. A checkout payment method fails for some customers on one mobile browser. What is the right route?

Put it into practice

Take your top ten backlog ideas and assign each to fix, just do it, test, research or drop, with a one-line rationale.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.