A/B testing: definition and how many visitors you need

A/B testing compares two versions of an element to find which converts better. Its difficulty is not technical, it is statistical.
3 min read
Believemy logo

Two headlines, two page versions, keep the winner: the principle looks unassailable. The problem is that most tests run by small operations measure nothing, and still produce a winner.

That false winner costs more than no test at all, because it drives a lasting decision off pure chance.


Definition

A/B testing means showing two versions of an element simultaneously to two randomly assigned groups of visitors, then comparing their Conversion rate.

The word "simultaneously" is a condition, not a detail. Comparing one week to the next also measures the day, the season and the traffic source.


The volume required, which surprises everyone

Does your test prove anything?
500
3 %
+0.6 pt
Visitors needed per version
13,912
You have 500 per version

At this volume that gap is indistinguishable from chance. A winner declared here would be a false winner.

The threshold is calculated for a two-sided test at 95 percent confidence and 80 percent power, the common convention. It does not replace a statistical tool, it gives the order of magnitude, which is nearly always enough to know whether to wait.

Set the visitor count and the gap between versions above. What appears is the one thing worth remembering: with a few hundred visitors, a two-point gap is indistinguishable from chance. Settling a realistic gap often takes several thousand visitors per version.

Warning

Direct consequence for a small business: on a page receiving three hundred visitors a month, A/B testing cannot prove anything. By the time you reach sufficient volume the season has changed and the test is invalid.


What to do without the volume

MethodWhat it gives
Narrated reading by five peopleThe sticking points, in an hour
A bold change measured over a quarterA gap large enough to be visible
A question asked of buyersThe real reason for the purchase
Watching where people leave the pageThe precise point of loss

The second row is the right strategy at low volume: do not test two shades, change boldly and measure over a long period. A ten-point gap shows even with little traffic; a one-point gap never will.


Frequently asked questions

Question

How long should a test run?

At minimum two full weeks, to cover every day of the week, and until the required volume is reached. Stopping a test as soon as one version pulls ahead is the most common error: gaps often reverse.


Question

Can I test several things at once?

Not at small volumes. Every extra variant divides traffic and multiplies the visitors needed. On low traffic, one change per test, and one test at a time.


Question

What should I test first?

Whatever everyone sees and sits close to the decision: the headline, the price, the wording of the offer. Testing an element seen by 5 percent of visitors takes twenty times the traffic for the same result.


Question

Is a losing test wasted time?

No, it is preserved information: you know that avenue returns nothing, and you will not revisit it in six months. Record results, negative ones included, or you will re-run the same tests.

Related terms

Discover our online business glossary

Every online business term explained in plain language: acquisition, recurring revenue, conversion, pricing, payments. Clear definitions and real numbers for founders and solopreneurs.

Share this article

Want to help us? Share this article on your networks or even better: on your site, in an article or in your newsletter.