Conversion Rate Optimization (CRO)Personalisation and running a CRO programme · Lesson 19 of 20
Running a CRO programme
Video lecture
Running a CRO programme
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Running a CRO programme
One test can give you a win. A programme gives you a compounding advantage. The difference isn't the tool you buy. It's roles, rhythm, metrics, a learning library and the way you handle stakeholders. In this lecture you'll learn how to move from one-off projects to a programme, who does what, a cadence that keeps work flowing, the programme metrics that matter, how to build a learning library your whole company can search, including with AI, and how to manage the highest paid person's opinion. Then you'll watch me query a learning library the right way.
0:42 Why a programme
Why a programme? Because one-off tests produce one-off learnings, and they're often forgotten. A programme runs research, hypotheses, tests and learning continuously, so each round is smarter than the last. It protects testing capacity from random requests. It creates institutional memory that survives staff changes. And it makes CRO legible to leadership, as a steady system for better decisions rather than a series of lucky headlines.
1:11 Roles
Roles. A CRO lead or strategist owns the roadmap, prioritisation and stakeholder alignment. A researcher or analyst owns analytics, qualitative research and test analysis. A designer and copywriter build variants. A developer builds and QAs them, including server-side changes. And stakeholders from product, brand, legal and customer service are consulted where relevant. In a small team, one person may hold three of these roles, and that's fine, as long as each responsibility is explicitly owned.
1:44 Cadence
Cadence. Weekly: check running tests for sample ratio mismatch, guardrails and QA issues, groom the backlog, and add new research inputs. Fortnightly: a prioritisation session, launch the next tests, share one learning. Monthly: a results review with stakeholders and a research sprint, like surveys, recordings or user tests. Quarterly: a programme review of velocity, win rate, learning rate, cumulative impact and roadmap themes. Rhythm beats heroics.
2:13 Programme metrics
Measure the programme, not just individual tests. Velocity: tests launched per month, where quality matters more than quantity. Win rate: the share of tests with a positive, significant result. Very high win rates can mean timid tests or flawed statistics; very low rates may mean weak research. Learning rate: tests that produced a documented, reusable insight, including losses. Cumulative estimated impact, conservatively estimated. And cycle time, from idea to decision.
2:43 The learning library
The learning library is the heart of a programme. Store every test, including losers and inconclusives, as a structured record: ID, dates, page, hypothesis, mechanism, variants with screenshots, the primary metric result with its interval, guardrails, the decision, learnings and tags. Tag by mechanism, like anxiety reduction, clarity, social proof or friction, so you can see patterns: anxiety reduction on product pages works well for us; urgency tests haven't. New team members learn in days, and nobody re-runs last year's test by accident.
3:19 Simple example: UK homeware library
A simple example. A small agency runs CRO for a UK homeware store. After six months, their library holds fourteen tests. Tagging by mechanism shows a clear pattern: tests that reduced delivery and returns anxiety won or were positive four times out of five, while tests that added urgency messaging were flat or negative. That single insight reshapes the next quarter's roadmap, and it's something no individual test could have told them.
3:50 Realistic example: Dubai agency quarter (illustrative)
Now a realistic scenario with illustrative details. A growth agency in Dubai runs CRO for a mid-sized ecommerce client with enough traffic for about two tests a month on key templates. Month one: a research sprint with an analytics audit, three hundred survey responses, six user tests and a heuristic review. They fix five bugs, implement three obvious improvements and build a backlog of forty ideas. Month two: they launch two tests, delivery estimates on product pages and wallet placement in checkout. Month three: one winner, one inconclusive, two follow-ups launched, and a quarterly review that leads with learnings and conservative impact estimates, not a single headline number.
4:37 Watch me do it: AI search over the library
Watch me query a learning library with AI, the right way. Our test records are exported as Markdown files. I attach them to my AI assistant and ask: what have we learned about reducing delivery anxiety on mobile product pages? But I add three rules. Answer only from the records. For each relevant test, give the ID, dates, change, result with its interval, decision and mechanism tag. And list the gaps, the questions we haven't tested yet. If the records don't answer the question, say so. The answer comes back with four test IDs I can open and check, and two untested ideas. That's a research briefing in two minutes, and every claim is traceable.
5:27 Stakeholders
Stakeholder management. Share learnings, not just wins, so losing tests are seen as valuable. Agree decision rules before tests launch, so nobody argues about what counts as a win afterwards. Involve legal, brand and customer service early for sensitive tests like pricing, claims and cancellation flows. And handle the HiPPO respectfully: turn opinions into hypotheses, score them like everything else, and let evidence decide.
5:55 Governance
Governance keeps a programme safe as it scales. Pre-register every test: hypothesis, metrics, sample size and decision rule. Put review gates in place for legal and compliance on sensitive tests, plus an accessibility and ethics checklist for every variant. Keep a calendar of running tests by page, to avoid collisions, and use mutually exclusive groups where needed. And practise tool hygiene: remove finished test code, and audit scripts every quarter for speed and privacy. Every script on your site has a cost.
6:31 The minimum toolkit
Which tools does a programme actually need? Fewer than you'd think. You need analytics, like GA4 or a product analytics tool. A behaviour and feedback tool, like Microsoft Clarity or Hotjar, plus on-site polls. A way to run user research, like Lyssna or Maze. An experimentation platform suited to your needs, whether marketer-led, feature flags, or both. Somewhere to keep the backlog and the learning library, like Notion, Airtable or Jira. And an AI assistant to help with synthesis and drafting. Pick the fewest tools that cover research, testing, analysis and knowledge. Every script on your site costs speed and carries privacy obligations, so audit them every quarter.
7:18 Maturity stages
Maturity comes in stages. Ad hoc: occasional tests driven by opinions. Emerging: regular tests and basic prioritisation. Structured: a consistent cadence, stakeholder reviews and a learning library. And embedded: experimentation is how decisions are made across teams, with holdouts, server-side testing and cross-team training. Most organisations move up one stage at a time. Buying an expensive platform rarely jumps you from ad hoc to embedded. Process and culture are the real constraints.
7:49 Common mistakes
Common mistakes. Measuring success by the number of tests alone. No learning library, so knowledge leaves with people. Running tests without stakeholder buy-in, then being overruled on rollout. Letting tests run forever because nobody owns the decision. Skipping documentation when busy. And tool-first thinking.
8:08 Recap
Recap. A programme turns tests into compounding learning. Define roles, run a weekly to quarterly cadence, measure velocity, win rate, learning rate, impact and cycle time, and keep a structured learning library that anyone, and any AI assistant, can search with traceable answers. Manage stakeholders with pre-agreed decision rules, and govern tests with pre-registration, review gates and tool hygiene. Try this now: set up a learning library with the record template from the lesson text, add your last five tests, and ask an AI assistant one question using the answer only from the records prompt.
From projects to a programme
One-off tests produce one-off learnings. A programme produces compounding gains because research, hypotheses, testing and learning run continuously with clear roles and rhythms.
Roles
| Role | Responsibilities | In a small team |
|---|---|---|
| CRO lead / strategist | Owns roadmap, prioritisation, stakeholder alignment | Often the marketer or founder |
| Researcher / analyst | Analytics, qualitative research, test analysis | Same person or freelancer |
| Designer / copywriter | Variant design and copy | Shared with marketing |
| Developer | Builds and QAs variants, server-side changes | Part-time or agency |
| Stakeholders | Product, brand, legal, customer service | Consulted on relevant tests |
Cadence
WEEKLY Test status check (SRM, guardrails, QA issues); backlog grooming; new research inputs
FORTNIGHTLY Prioritisation session; launch next tests; share one learning
MONTHLY Results review with stakeholders; research sprint (surveys, recordings, user tests)
QUARTERLY Programme review: velocity, win rate, cumulative impact, roadmap themesProgramme metrics
Measure the programme, not just individual tests:
- Velocity — tests launched per month (quality matters more than quantity).
- Win rate — share of tests with a positive, significant result. Very high win rates can indicate timid tests or flawed statistics; very low rates may indicate weak research.
- Learning rate — tests that produced a documented, reusable insight (including losers).
- Cumulative estimated impact — conservatively estimated, with caveats.
- Cycle time — from idea to decision.
The learning library
Store every test in a searchable repository:
Test ID | Name | Dates | Page/template | Hypothesis | Mechanism | Variants (screenshots)
Primary metric result + interval | Guardrails | Decision | Learnings | Tags (e.g. anxiety, delivery, pricing)Tags by mechanism (anxiety reduction, clarity, social proof, friction) let you see patterns: "anxiety-reduction tests on product pages have performed well; urgency tests have not". New team members learn quickly, and you avoid re-running old tests.
Tools
- Analytics (GA4 or similar) for funnels and segments.
- Behavioural tools for heatmaps and recordings.
- Survey and user-testing tools.
- Experimentation platform — client-side, server-side or feature-flag based. Choose based on traffic, technical resources, performance impact and statistical approach.
- Project and knowledge tools for backlog and learning library.
Stakeholder management
- Share learnings, not just wins; a learning culture tolerates losing tests.
- Pre-agree decision rules (what counts as a win) before tests launch.
- Involve legal, brand and customer service early for sensitive tests (pricing, claims, cancellation flows).
- Handle the HiPPO respectfully: turn opinions into hypotheses and let evidence decide.
Worked example: a small agency's CRO retainer
A growth agency in Dubai runs CRO for a mid-sized e-commerce client with enough traffic for about two concurrent tests per month on key templates. Their first quarter:
Month 1: Research sprint (analytics audit, 300 survey responses, 6 user tests, heuristic review)
Fixed 5 bugs; implemented 3 obvious improvements; built backlog of 40 ideas
Month 2: Launched 2 tests (delivery estimate on PDP; checkout wallet placement)
Month 3: Analysed; 1 winner, 1 inconclusive; launched 2 follow-ups; quarterly reviewThe quarterly report leads with learnings and conservative impact estimates, not with a single headline uplift.
Tooling map for a 2026 CRO programme
| Need | Examples (not endorsements) |
|---|---|
| Analytics | GA4, Adobe Analytics, Amplitude, Mixpanel, PostHog, Piwik PRO, Matomo |
| Behaviour and feedback | Microsoft Clarity, Hotjar/Contentsquare, FullStory, Mouseflow; on-site polls |
| UX research | Lyssna, Maze, UserTesting, Lookback, Optimal Workshop |
| Experimentation | VWO, Optimizely, AB Tasty, Kameleoon, Convert; GrowthBook, PostHog, Statsig, LaunchDarkly |
| Data | BigQuery / Snowflake / Databricks for warehouse-native analysis |
| Backlog and learning library | Notion, Airtable, Jira, Confluence, Google Sheets |
| AI assistants | Claude, ChatGPT, Gemini, Copilot for research synthesis, drafting and analysis support |
Pick the fewest tools that cover research, testing, analysis and knowledge. Every script on your site has a speed and privacy cost.
Hands-on: make the learning library AI-searchable
A learning library becomes more valuable when anyone can ask it questions ("What have we learned about delivery messaging on mobile?"). Keep each test as a structured record (the template above), export it as Markdown or CSV, and use an AI assistant or your workspace's built-in AI search over those records:
You are helping me search our experiment learning library (attached records).
Question: What have we learned about reducing delivery anxiety on mobile product pages?
Answer only from the records. For each relevant test give: ID, dates, change, result (effect + interval),
decision, and mechanism tag. Then list gaps: questions we have not tested yet.
If the records do not answer the question, say so."Answer only from the records" and asking for test IDs make it easy to verify every claim.
Governance essentials
- Pre-registration: hypothesis, primary metric, guardrails, sample size and decision rule recorded before launch.
- Review gates: legal/compliance for pricing, subscriptions and claims; accessibility and ethics checklist for every variant.
- Collision management: a calendar of running tests by page/template; mutually exclusive groups where needed.
- Tool hygiene: remove finished test code; audit scripts quarterly for speed and privacy.
Common mistakes
- Measuring success by number of tests alone.
- No learning library, so knowledge leaves with people.
- Running tests without stakeholder buy-in, then being overruled on rollout.
- Letting tests run indefinitely because nobody owns the decision to stop and analyse them.
- Skipping post-test documentation when the team is busy — the learning is then lost for good.
- Tool-first thinking: buying a platform before having research and hypotheses.
Programme maturity stages
| Stage | Characteristics | Next step |
|---|---|---|
| Ad hoc | Occasional tests driven by opinions | Establish a research log and backlog |
| Emerging | Regular tests, basic prioritisation | Add test plans, SRM checks and a learning library |
| Structured | Consistent cadence, stakeholder reviews | Add programme metrics and mechanism tagging |
| Embedded | Experimentation is how decisions are made across teams | Holdouts, server-side testing, cross-team training |
Most organisations move up one stage at a time. Trying to jump from ad hoc to embedded by buying an expensive platform rarely works; process and culture are the real constraints.
Key takeaways
- A programme compounds gains through continuous research, testing and learning.
- Track velocity, win rate, learning rate, cycle time and conservative impact.
- A tagged learning library prevents repeated tests and spreads knowledge.
- Agree decision rules with stakeholders before tests launch.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Set up a learning library template and log at least three past changes or tests with mechanism tags and learnings.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.