Ship it or not? An A/B test game
Your team has redesigned the checkout page. Today 5% of visitors buy. In each round, a coin flip decides in secret whether the new version B is truly better (it lifts sales by 15%) or exactly as good as the old version A.
Nobody is adding noise on purpose. Each visitor simply buys or doesn’t, at random, so even two identical pages rarely show the same conversion rate. How big that wobble is depends only on how many visitors you have: with 2,000 per version, pure chance can make B look up to about 27% better or worse than A, which easily hides a real 15% lift. With 15,000 per version, the wobble shrinks to about ±10%. The line under the slider shows this range as you move it.
Pick how many visitors to send to each version, press start, and watch the evidence come in. You can call it at any moment: Ship B or Keep A. Then the truth is revealed.
What to try
- Run a few rounds at 2,000 visitors. When B is truly better, the test usually misses it. A small test is not “cheap”, it is mostly noise. Slide up to around 15,000 and the power figure passes 80%, the level most teams aim for.
- Watch the blue line before the test ends. Even when B is a dud, the p-value wanders and often dips below 0.05 for a moment. If you ship whenever it first looks significant, you’ll ship far more duds than the 5% the textbook promises. This is the peeking problem, and it is why teams fix the sample size up front or use sequential methods built for early stopping.
- Notice how small the real effect is. A 15% lift here is 5% → 5.75%. Most real product changes are this small or smaller, which is why good experiments need patience more than clever dashboards.
In product work many launch decisions end in a readout like this one. The hard part is rarely the statistics; it is deciding in advance what would change your mind, and then sticking to it.
Comments
Comments load here, powered by GitHub Discussions.