Newsletter

One email on Fridays, and nothing else.

  • Practical B2B tips

  • 4-min read on Fridays

  • For anyone in B2B growth

How to apply

Start with clear hypotheses. Rather than vague 'test if this email performs better,' write: 'We believe that subject lines addressing specific ROI metrics will increase open rates by 5% because our target audience evaluates on financial impact.' Specific hypotheses make success criteria clear.

Change one variable per test. Testing multiple changes simultaneously makes it impossible to identify which change drove results. If you test both subject line and send time, and performance improves, which element caused it? Single-variable tests provide clear attribution.

Define sample size and duration before starting tests. Decide how many test recipients you need and how long you'll run tests before analysing results. This prevents temptation to stop tests early when results look positive (which often leads to false positives).

Analyse statistical significance, not just percentage difference. A 5% improvement might be meaningful or noise depending on sample size. Tools like A/B test calculators show whether improvements are statistically significant. Only implement changes where improvements are statistically significant at 95% confidence level.

Hypothesis testing means you decide what change you're making and what you expect to happen before you make it, then let the data tell you whether you were right. Instead of shipping a new headline because it feels better and assuming it worked, you write down a clear prediction, change one thing, measure the result, and keep or bin the change based on what actually happened.

The point is to learn, not to guess. Founders and teams are full of hunches: 'this subject line will land better', 'this layout will convert'. Hunches are wrong more often than we'd like to admit. A proper test replaces the loudest opinion in the room with a number, which saves you from rolling out a worse experience to everyone.

A test has a few moving parts. A hypothesis: a specific prediction and the reason behind it. One variable: the single thing you're changing, so you can attribute the result. A control (the current version) and a treatment (the new version). A sample size big enough that the result isn't just noise. A duration long enough that a quiet Tuesday or a busy month-end doesn't skew it. And a success metric you agreed on up front. B2B usually needs bigger samples than B2C because the volumes are lower, so a meaningful email test might need 500 recipients per side rather than 50, plan the scope for that.

Why it matters

Testing first stops you shipping expensive mistakes to your whole base. New pricing, new messaging, a redesigned onboarding, get those wrong at full scale and you've burned trust and revenue. Prove it on a slice first.

It also compounds. Every test teaches you something about what your customers respond to, and those lessons stack up: which angles convert, which features people actually use, which offers move the needle. Over a year you're not guessing any more, you're deciding from a library of what's worked.

And it kills bike-shedding. When a test settles the argument, the team stops debating opinions and starts shipping. Decisions get faster and less political because they're grounded in data, not whoever argues hardest.

How to apply it

Write a real hypothesis. Not 'test if this email does better', but 'subject lines naming a specific ROI figure will lift open rates by 5%, because our buyers evaluate on financial impact'. The prediction makes the success bar obvious before you start.

Change one thing. If you swap the subject line and the send time and opens go up, you've learned nothing about which one did it. One variable per test buys you clean attribution.

Set sample size and duration before you start, and don't peek. Decide how many people and how many days up front. Stopping a test early because it looks good is the single fastest way to fool yourself with a false positive.

Judge significance, not the raw percentage. A 5% lift is either real or noise depending on how many people saw it. Check it hits 95% confidence before you roll anything out.

Examples

Testing an onboarding sequence. Say you're running activation experiments with Amplitude. Your hypothesis is that new users stall because they don't grasp the value, not the mechanics. You ship two onboarding flows, a plain feature walkthrough versus a value-first version (three use-cases, then the matching features), and tag both cohorts in Amplitude so you can watch each group's path to full engagement over 30 days. The value-first group activates at 45% against 28% for the walkthrough, a 17-point gap that clears significance. You roll the new flow out to everyone and overall activation climbs from 28% to 42%.

Testing an email angle. Say you're sending a campaign through Brevo. You predict a cost-reduction subject line will beat an efficiency one, because the CFO signs off and the CFO cares about cost. You split 5,000 contacts per side in Brevo, hold it for seven days, and don't touch it. Cost-reduction lifts opens from 18% to 24% and clicks from 4% to 6.2%, both significant. You move every sales follow-up to lead with cost, and the downstream conversion follows.

Testing a landing page layout. Say you're running the test in VWO. Your hypothesis: putting customer logos above the fold lifts form submissions by 10%, because the social proof calms evaluation nerves. You split traffic 50/50 in VWO, 2,000 visitors per version over a week. Submissions go from 3.2% (logos below the fold) to 5.1% (logos above it), a 59% jump that's statistically significant. You push the change to every landing page and lift conversion across the board.

Email campaign testing messaging angle

An enterprise software company tested email subject lines. Hypothesis: Addressing cost reduction (ROI angle) would outperform process improvement (efficiency angle) because CFOs (key decision-maker) prioritise cost. Test group 1 received emails with cost-reduction messaging. Test group 2 (control) received standard messaging. Sample size: 5000 per group. Run duration: 7 days. Cost-reduction messaging improved open rate from 18% to 24% and click rate from 4% to 6.2%. Improvements were statistically significant. The company shifted all sales follow-up emails to emphasise cost reduction, improving conversion rates downstream.

SaaS testing onboarding sequence

A SaaS company hypothesised that new users struggling to complete onboarding were due to not understanding feature value. They tested two versions: (1) standard onboarding (feature walkthrough), (2) value-focused onboarding (showing three use cases, then walking through corresponding features). Both groups were tracked over 30 days. Users in the value-focused group (group 2) reached full product engagement at 45% rate, versus 28% for the standard group. This 17-point improvement was statistically significant. The company rolled out value-focused onboarding to all new users, improving overall user activation rate from 28% to 42%.

Why it matters

Hypothesis testing prevents expensive mistakes. Implementing untested changes across all customers risks poor outcomes. Testing first allows you to confirm changes drive desired results before full implementation. This prevents launching ineffective messaging, pricing, or features to the entire customer base.

Hypothesis testing compounds learning over time. Each test provides insights into what resonates with your customers. Accumulated test results reveal patterns: which headlines work, which offers convert, which features drive engagement. These patterns guide future decisions with increased confidence.

Hypothesis testing improves team efficiency. Rather than debating whether a change will work, teams run tests and let data decide. This reduces bike-shedding, accelerates decision making, and builds team confidence in decisions: they're based on data, not politics or strongest opinion.

Agency testing landing page layout

A B2B agency tested landing page layouts. Hypothesis: Placing customer logos prominently above the fold would increase form submissions by 10% because social proof reduces evaluation anxiety. They split traffic 50/50: version A (logos below fold), version B (logos above fold, prominent). 2000 visitors per version, one-week duration. Form submission rate: version A (3.2%), version B (5.1%). This 1.9-point improvement (59% increase) was statistically significant. The agency updated all landing pages to feature customer logos prominently, systematically improving conversion rates across campaigns.

Articles

  • Article

    Statistical significance is just the beginning. Learn how to interpret results correctly, avoid false positives, and turn winning experiments into permanent improvements across your growth engines.

  • Article

    Most experiments fail before they start because the hypothesis is vague or untestable. Learn how to write hypotheses that are specific enough to prove or disprove and tied to metrics that matter.

  • Article

    A folder full of interview notes is worthless if nothing changes. Learn how to spot patterns across conversations and turn what you heard into better copy, sharper ads, and stronger sales conversations.

  • Article

    A winning test means nothing if the setup was flawed. Learn how to configure experiments properly in VWO, ad platforms, and email tools so your results are actually valid.

  • Article

    Random testing wastes time and teaches you nothing. Learn how to collect experiment ideas systematically and prioritise them based on potential impact so you always know what to run next.

  • Article

    Build a knowledge base from past experiments so new tests build on proven insights instead of starting from scratch every time.

All 20 articles under A/B testing and experimentation
FAQ

Questions about this topic

Academy

Growth Academy

Start free

A free account opens the first course and keeps your progress.

  • A free course

  • Track your own skills

  • Every playbook you unlock