Article

Building your backlog

Random testing wastes time and teaches you nothing. Learn how to collect experiment ideas systematically and prioritise them based on potential impact so you always know what to run next.

Updated 1 September 2026

Article

Random testing wastes time and teaches you nothing. Learn how to collect experiment ideas systematically and prioritise them based on potential impact so you always know what to run next.

Newsletter

One email on Fridays, and nothing else.

  • Practical B2B tips

  • 4-min read on Fridays

  • For anyone in B2B growth

Summary

This lesson walks you through building an experiment backlog from scratch. You will learn where good experiment ideas come from, how to capture them in a simple format, and how to score them so you always know which test to run next.

A backlog is just a list of experiment ideas, scored and sorted so you always know what to test next. Without one, you end up testing whatever someone suggested in the last meeting. With one, you can make deliberate choices about where to focus.

The format doesn't matter much. I've used Notion, spreadsheets, and dedicated tools over the years. What matters is that you can add ideas quickly, score them consistently, and sort by priority. Everything else is decoration.

The real work isn't maintaining the backlog. It's generating ideas worth testing and being honest about which ones will actually move revenue.

Good experiment ideas rarely appear out of nowhere. They come from two sources: anomalies in your data and insights from customer research.

Data anomalies are the starting point. You notice that a landing page has high traffic but low conversion. You see that one email in a sequence has a 40% drop-off. You spot that mobile users behave completely differently from desktop users. These patterns tell you where something is broken or underperforming, but they don't tell you why.

Customer research fills in the why. You interview users who abandoned a form and discover the pricing section confused them. You talk to customers who converted quickly and learn they almost left because they couldn't find a specific feature. You review session recordings and see people scrolling past your call-to-action without noticing it.

The combination is what generates testable ideas. Data shows you where to look. Research tells you what might be wrong. Together, they give you a hypothesis worth testing.

Other sources can feed the backlog too: competitor analysis, team brainstorms, support tickets, sales objections. But treat these as secondary. An idea that came from "I saw a competitor do this" is weaker than an idea that came from "our data shows a problem and our research suggests a cause."

Every prioritisation framework is essentially the same. You're trying to estimate impact, confidence, and effort, then sort by some combination of those factors.

ICE scores each idea from 1-10 on Impact, Confidence, and Ease, then averages them. PIE uses Potential, Importance, and Ease. Some teams use weighted formulas. Others just use high/medium/low ratings.

Pick whatever framework your team will actually use. The specific scoring method matters less than being consistent and honest. Most teams inflate scores for ideas they're excited about and deflate scores for ideas that seem boring but important. Fight that tendency.

The only filter that really matters is revenue impact. An experiment that improves a metric nobody cares about is a waste of time, even if it wins. Before scoring any idea, ask: if this works, how does it affect revenue? If the answer is unclear or indirect, score it lower.

A backlog that never gets reviewed becomes a graveyard of abandoned ideas. A backlog that gets reviewed constantly becomes a distraction from actually running tests.

Review your backlog when you need to decide what to test next. That might be weekly if you're running fast experiments on a high-traffic site, or monthly if you're in a slower B2B context with longer test cycles.

During each review, do three things. First, add any new ideas that came up since the last review. Second, re-score ideas if you've learned something that changes your estimate of impact or confidence. Third, archive ideas that are no longer relevant because the page changed, the problem was solved another way, or you've learned enough to know the idea won't work.

The goal is a backlog that's small enough to be useful. If you have 200 ideas sitting there, you're not going to read through them every time you need to pick a test. Aim for 20-30 active ideas, with older or lower-priority items archived somewhere you can search if needed.

Your backlog should reflect your current growth priorities, not just random opportunities to improve things.

If your quarterly focus is improving activation rate, your top-scored experiments should target that metric. If you're trying to increase deal size, your backlog should include tests on pricing pages and upgrade flows. The backlog isn't separate from your growth strategy; it's one of the tools for executing it.

This also means your priorities will shift. An experiment idea that scored highly last quarter might drop in priority this quarter because you're focused on a different engine. That's fine. Re-score based on current priorities, not historical scores.

When you review your backlog, start by reminding yourself what you're trying to improve right now. Then look at which ideas directly target that metric. Those go to the top.

A backlog is a simple thing: a list of ideas, scored and sorted. But maintaining one consistently is what separates teams that learn from experimentation from teams that just run random tests.

Start by documenting where your ideas come from. Score them honestly based on revenue impact. Review regularly but not obsessively. And keep the list connected to whatever growth priority you're focused on this quarter.

The backlog itself won't make you better at experimentation. But it will make sure you're always testing the most important thing, not just the most recent suggestion.

More articles

  • Article

    Statistical significance is just the beginning. Learn how to interpret results correctly, avoid false positives, and turn winning experiments into permanent improvements across your growth engines.

  • Article

    Most experiments fail before they start because the hypothesis is vague or untestable. Learn how to write hypotheses that are specific enough to prove or disprove and tied to metrics that matter.

  • Article

    A folder full of interview notes is worthless if nothing changes. Learn how to spot patterns across conversations and turn what you heard into better copy, sharper ads, and stronger sales conversations.

  • Article

    A winning test means nothing if the setup was flawed. Learn how to configure experiments properly in VWO, ad platforms, and email tools so your results are actually valid.

  • Article

    Build a knowledge base from past experiments so new tests build on proven insights instead of starting from scratch every time.

  • Article

    Take validated wins from one channel and systematically test whether they work in others to multiply the impact of what already works.

All 19 articles under A/B testing and experimentation
FAQ

Questions about this topic

Academy

Growth Academy

Start free

A free account opens the first course and keeps your progress.

  • A free course

  • Track your own skills

  • Every playbook you unlock