Email subject line testing at a SaaS company

A B2B SaaS firm testing email subject lines to their database of 50,000 prospects found that 'Your team is losing £5,000 per week on X' significantly outperformed 'Learn how to improve X efficiency' (32% open rate vs 18%). The test was run to 5,000 people, well above the sample size needed for significance. The winning line was then used in all future campaigns to that segment.

Call-to-action button colour test in a payment platform

A fintech company tested CTA button colours on their checkout page. The contrasting orange button outperformed the muted grey button by 4.3%, a seemingly small lift. Applied across 2 million annual transactions, this translated to an additional £180,000 in revenue annually with zero product changes.

A/B testing is how you settle an argument with data instead of opinion. You take one thing , a page, an email, a button , make two versions, show version A to half your audience and version B to the other half, and see which one wins on a number you care about. Version A is the control (what you have now); version B is the variant (the change you're trying). Because the only thing different between the two groups is that one change, any difference in results is down to that change , not luck, not the weather, not which half got the keener visitors.

The discipline is changing ONE thing at a time. If you rewrite the headline AND the button AND the form fields all at once and conversions jump 15%, you've learned nothing , you don't know which change did the work. Change one element, and a 15% lift actually tells you something you can reuse.

The other half of the discipline is waiting for enough data. A 20% lift off 40 visitors is noise; the same lift off 4,000 visitors is a signal you can bank. Most tools won't even call a winner until each version has had a few hundred conversions, and they're right to make you wait.

It matters because the wins stack. A few percent on your email open rate, a few more on your landing-page conversion, a bit on checkout , each one is small on its own, but they compound down the funnel and across the year into real money.

Examples

Say you're running a landing page for a new offer and you genuinely can't decide between two headlines. Instead of arguing about it, you build both as variants in Unbounce, split your paid traffic evenly, and let it run until each version has a few hundred conversions. "Close your strategy gaps in 90 days" beats "Strategy consulting for SaaS" at 8.2% versus 6.1% , the specific, time-bound promise wins, and now you know to write the next page that way too.

Say you're sending a campaign to 50,000 prospects and you want to know which subject line lands. In Brevo you set up a subject-line test that sends both versions to a small slice first, measures the open rate, then automatically sends the winner to everyone else. "Your team is losing 5,000 GBP a week on X" pulls a 32% open rate against 18% for "Learn how to improve X efficiency" , the loss-framed line wins, and it becomes your template for that segment.

The trap with testing buttons and layouts is that you can't always tell why a version won. Pair the test with a session-replay tool like Microsoft Clarity: when your contrasting orange checkout button beats the muted grey one by 4.3%, the heatmaps and recordings show you people were genuinely missing the old button, not just clicking faster. That 4.3% sounds tiny until you apply it across two million transactions a year , then it's six figures of revenue from a colour change and zero product work.

Landing page headline testing for a consulting firm

A management consulting firm tested two headlines: 'Strategy consulting for growth-stage SaaS' vs 'Close your critical strategy gaps in 90 days'. The second variant converted at 8.2% compared to 6.1%. The specificity of the outcome ('close gaps') and the time constraint ('90 days') resonated more than the generic category positioning.

How to apply

Define your hypothesis and metric

Start by identifying one element to test and the outcome you're measuring. Don't test 'everything looks better'—test 'changing the button from blue to green will increase form completions by 5%'. The metric must be trackable: conversion rate, click-through rate, email open rate, or time on page.

Split your audience randomly

Divide your traffic or user base equally between control and variant. Randomisation prevents selection bias. If your high-intent users all see version B, you can't claim B is better—it's just attracting higher-intent visitors.

Run the test long enough

Reach statistical significance before concluding. A 10% lift from 50 clicks is noise. A 10% lift from 5000 clicks is signal. Most platforms require at least 100-200 conversions per variant before results are reliable.

Document and iterate

Record every test, winner, and insight. This creates institutional memory and prevents repeating failed experiments. Winning tests often become the baseline for the next test—continuous improvement compounds.

Why it matters

Prevents expensive guesses

Without testing, marketing decisions rest on opinion, trends, or what competitors do. A redesigned homepage might look beautiful but underperform. An email with personalisation might get lower engagement than expected. Testing removes the guesswork and validates assumptions before rolling out changes across your entire audience.

Builds evidence for larger decisions

A single A/B test might improve conversion by 2%. But when you run dozens of tests across your funnel, each small improvement multiplies. The tests also generate internal credibility—stakeholders see the data and buy in to further optimisation work. This accelerates decision-making across product, design, and marketing.

Uncovers unexpected insights

A/B tests often reveal counter-intuitive results. The longer form might outperform the short one. The urgent copy might underperform the educational copy. Testing exposes what your actual audience responds to, not what you assumed they would.

Articles

  • Article

    Statistical significance is just the beginning. Learn how to interpret results correctly, avoid false positives, and turn winning experiments into permanent improvements across your growth engines.

  • Article

    Most experiments fail before they start because the hypothesis is vague or untestable. Learn how to write hypotheses that are specific enough to prove or disprove and tied to metrics that matter.

  • Article

    A folder full of interview notes is worthless if nothing changes. Learn how to spot patterns across conversations and turn what you heard into better copy, sharper ads, and stronger sales conversations.

  • Article

    A winning test means nothing if the setup was flawed. Learn how to configure experiments properly in VWO, ad platforms, and email tools so your results are actually valid.

  • Article

    Random testing wastes time and teaches you nothing. Learn how to collect experiment ideas systematically and prioritise them based on potential impact so you always know what to run next.

  • Article

    Build a knowledge base from past experiments so new tests build on proven insights instead of starting from scratch every time.

All 20 articles under A/B testing and experimentation
FAQ

Questions about this topic

Academy

Growth Academy

Start free

A free account opens the first course and keeps your progress.

  • A free course

  • Track your own skills

  • Every playbook you unlock