Skip to content

Answer

How much traffic do you need for A/B testing?

By CartKernel ยท Last reviewed

In short

It depends on two numbers you already have: your current conversion rate and the smallest change you would act on. A low baseline rate and a small change to detect need an enormous sample, while a higher baseline and a bold change need far less. Use the standard approximation to calculate it before you build anything. Most stores discover they can only test large changes, which is a useful finding, because small refinements were never going to be measurable at their volume anyway.

The sample size comes out of an equation, not a rule of thumb

For a straightforward comparison of two versions, a widely used approximation puts the visitors needed per variant at sixteen times the baseline rate multiplied by one minus the baseline rate, divided by the square of the absolute difference you want to detect. That gives roughly eighty percent power at the usual significance level.

The important part is the square in the denominator. Halving the change you want to detect quadruples the traffic you need. This is why a store can measure a redesign of a product page and cannot measure a new button colour, and why an ambitious test is often the only kind worth running at moderate volume.

Be careful to express the difference in absolute terms. Moving from two percent to two and a half percent is a half point difference in absolute terms, and putting the relative figure into the formula understates the requirement badly.

Run the calculation before designing the test. If the answer is more visitors than you will see this quarter, you have learned something useful before spending any development time.

Duration matters as much as sample size

Reaching the visitor count quickly does not make a test valid. Buying behaviour varies by day of week, by payday, by weather and by campaign schedule, so a test that runs for four days measures those four days rather than your store.

Run whole weeks, and at least two of them. That covers the weekday and weekend pattern and gives the traffic mix time to look like normal traffic rather than like whatever campaign happened to be live.

Allow for the gap between visit and purchase. In categories where people consider for several days, a visitor who saw the test on day one may buy on day nine, and stopping the test on day seven attributes their eventual order to nobody.

Avoid running a test across an event that changes behaviour. A sale, a stock-out on a hero product, a public holiday or a large campaign launch will move the numbers more than the change you are testing, and no amount of statistics will separate them afterwards.

What to do when you do not have the traffic

Test bigger things. A store that cannot detect a small refinement can often detect a wholesale change to a page: a different structure, a different offer, a different set of information. Those are also the changes most likely to matter.

Move up the funnel. Add to cart happens far more often than purchase, so a test measured on add-to-cart rate needs far less traffic to read. Treat it as a proxy, confirm afterwards that orders moved too, and be alert to changes that raise adds and lower purchases.

Use before-and-after measurement carefully where a split test is impossible. It is weaker because the world changes between periods, and it is not worthless if you segment by traffic source, compare like periods and avoid confounding events.

And spend the effort on research instead of on statistics. Session recordings, on-site surveys, support tickets and watching five real people use the checkout will tell a low-traffic store more in a week than an underpowered test will tell it in a quarter.

The mistakes that make a result meaningless

Stopping when the result looks good is the biggest one. Checking daily and ending the test the moment significance appears converts a rigorous method into a search for a favourable moment, and it produces winners that do not repeat.

Decide the sample size and the end date before you start, then look at the result once. If you need to monitor a test in flight, monitor it for breakage rather than for outcome.

Check that traffic actually split as intended. A large imbalance between the groups points at a technical problem with the assignment, and any result from an unbalanced test should be discarded rather than interpreted.

Run one test at a time on the same journey. Two overlapping tests on the product page and the cart interact, and the reported results for each will include the effect of the other. And record what you tested and what happened, including the failures, because a store's own testing history is the most valuable evidence it will ever have.

Sample size for one product page test

Baseline conversion rate
2.0 percent
Change worth acting on
2.0 to 2.4 percent
Absolute difference
0.4 percentage points
Visitors needed per variant
About 19,600
Visitors needed in total
About 39,200
At 6,000 relevant sessions a week
Around seven weeks
If only a smaller change were expected
Four times as long again

Illustrative figures using the standard approximation at conventional power and significance. The last row is the one that changes plans, because detection cost rises with the square of the precision you want.

Related questions

Can I test on a store with a few hundred sessions a day?

You can test large changes on high-frequency events such as add to cart, and you should not expect to measure small refinements on purchases. At that volume, qualitative research and fixing known problems will move revenue faster than a testing programme will.

Is it acceptable to stop a test early if one version is clearly ahead?

No, unless you planned for interim checks with a method that accounts for them. Looking repeatedly and stopping at the first favourable moment inflates the chance of calling a false winner, and those winners tend to disappear when the change is rolled out to everyone.

Should I test on sessions or on users?

Assign by user, so the same person sees the same version across visits, and measure the outcome consistently against whichever denominator you chose. Mixing them, by assigning per session and measuring per user, produces numbers that cannot be interpreted.

Find the leak.

A free Growth Analysis ranks what your store should fix first, by revenue at stake.