Article

Run it long enough to trust it, and never peek to a stop

Newsletter

One email on Fridays, and nothing else.

  • Practical B2B tips

  • 4-min read on Fridays

  • For anyone in B2B growth

Run it long enough to trust it, and never peek to a stop

Run it long enough to trust it, and never peek to a stop

Experiments fail honesty in two opposite ways, and both come from impatience with time. The first is stopping too early because the number looks good. The second is reading the number every day and stopping the moment it crosses the line. Both produce winners that are not real, and both are seductive precisely because they feel like decisiveness.

Conversion data is noisy. On day two, almost any variant can be ahead by chance, and if you stop the instant it looks like a winner, you will ship a coin flip and call it a strategy. The fix is to decide the duration and sample size before launch and then leave the test alone until it reaches them. The run length is a commitment, not a suggestion you revisit when the early numbers tempt you.

Peeking is the subtler trap because it masquerades as diligence. If you check a running test repeatedly and stop as soon as it shows significance, you have quietly run dozens of chances for noise to cross the line, and noise eventually will. The result is a false win you will struggle to reproduce. Look at the dashboard if you must, but the stopping rule was set in advance and the daily number does not get a vote.

Time also matters because buyers are not uniform across a week. Weekday traffic behaves differently from weekend traffic, the start of the month differs from the end, a campaign spike differs from a quiet stretch. A test that runs for two days samples a slice of behaviour and generalises from it badly. Running across at least one full weekly cycle, often two, is what makes the read represent your actual audience rather than Tuesday's.

The reward for patience is a result you can act on without flinching. A test that ran its committed length, hit its pre-set threshold, and held across a full cycle is a result you can build the next quarter on. A test you stopped early because it looked good is a result you will quietly distrust, and distrust is expensive: it makes you re-test things you already answered.

INTERVIEW EWOUD: Share a case where running a test longer changed the outcome, the early read pointed one way and the full run pointed another. How long did it run, and what would you have shipped if you had stopped on the early number?

More articles

  • Article

    Statistical significance is just the beginning. Learn how to interpret results correctly, avoid false positives, and turn winning experiments into permanent improvements across your growth engines.

  • Article

    Most experiments fail before they start because the hypothesis is vague or untestable. Learn how to write hypotheses that are specific enough to prove or disprove and tied to metrics that matter.

  • Article

    A folder full of interview notes is worthless if nothing changes. Learn how to spot patterns across conversations and turn what you heard into better copy, sharper ads, and stronger sales conversations.

  • Article

    A winning test means nothing if the setup was flawed. Learn how to configure experiments properly in VWO, ad platforms, and email tools so your results are actually valid.

  • Article

    Random testing wastes time and teaches you nothing. Learn how to collect experiment ideas systematically and prioritise them based on potential impact so you always know what to run next.

  • Article

    Build a knowledge base from past experiments so new tests build on proven insights instead of starting from scratch every time.

All 19 articles under A/B testing and experimentation
FAQ

Questions about this topic

Academy

Growth Academy

Start free

A free account opens the first course and keeps your progress.

  • A free course

  • Track your own skills

  • Every playbook you unlock