Run it long enough to trust it, and never peek to a stop
Run it long enough to trust it, and never peek to a stop
Experiments fail honesty in two opposite ways, and both come from impatience with time. The first is stopping too early because the number looks good. The second is reading the number every day and stopping the moment it crosses the line. Both produce winners that are not real, and both are seductive precisely because they feel like decisiveness.
Conversion data is noisy. On day two, almost any variant can be ahead by chance, and if you stop the instant it looks like a winner, you will ship a coin flip and call it a strategy. The fix is to decide the duration and sample size before launch and then leave the test alone until it reaches them. The run length is a commitment, not a suggestion you revisit when the early numbers tempt you.
Peeking is the subtler trap because it masquerades as diligence. If you check a running test repeatedly and stop as soon as it shows significance, you have quietly run dozens of chances for noise to cross the line, and noise eventually will. The result is a false win you will struggle to reproduce. Look at the dashboard if you must, but the stopping rule was set in advance and the daily number does not get a vote.
Time also matters because buyers are not uniform across a week. Weekday traffic behaves differently from weekend traffic, the start of the month differs from the end, a campaign spike differs from a quiet stretch. A test that runs for two days samples a slice of behaviour and generalises from it badly. Running across at least one full weekly cycle, often two, is what makes the read represent your actual audience rather than Tuesday's.
The reward for patience is a result you can act on without flinching. A test that ran its committed length, hit its pre-set threshold, and held across a full cycle is a result you can build the next quarter on. A test you stopped early because it looked good is a result you will quietly distrust, and distrust is expensive: it makes you re-test things you already answered.
INTERVIEW EWOUD: Share a case where running a test longer changed the outcome, the early read pointed one way and the full run pointed another. How long did it run, and what would you have shipped if you had stopped on the early number?