Two headlines, two page versions, keep the winner: the principle looks unassailable. The problem is that most tests run by small operations measure nothing, and still produce a winner.
That false winner costs more than no test at all, because it drives a lasting decision off pure chance.
Definition
A/B testing means showing two versions of an element simultaneously to two randomly assigned groups of visitors, then comparing their Conversion rate.
The word "simultaneously" is a condition, not a detail. Comparing one week to the next also measures the day, the season and the traffic source.
The volume required, which surprises everyone
At this volume that gap is indistinguishable from chance. A winner declared here would be a false winner.
The threshold is calculated for a two-sided test at 95 percent confidence and 80 percent power, the common convention. It does not replace a statistical tool, it gives the order of magnitude, which is nearly always enough to know whether to wait.
Set the visitor count and the gap between versions above. What appears is the one thing worth remembering: with a few hundred visitors, a two-point gap is indistinguishable from chance. Settling a realistic gap often takes several thousand visitors per version.
Direct consequence for a small business: on a page receiving three hundred visitors a month, A/B testing cannot prove anything. By the time you reach sufficient volume the season has changed and the test is invalid.
What to do without the volume
| Method | What it gives |
|---|---|
| Narrated reading by five people | The sticking points, in an hour |
| A bold change measured over a quarter | A gap large enough to be visible |
| A question asked of buyers | The real reason for the purchase |
| Watching where people leave the page | The precise point of loss |
The second row is the right strategy at low volume: do not test two shades, change boldly and measure over a long period. A ten-point gap shows even with little traffic; a one-point gap never will.
Frequently asked questions
How long should a test run?
At minimum two full weeks, to cover every day of the week, and until the required volume is reached. Stopping a test as soon as one version pulls ahead is the most common error: gaps often reverse.
Can I test several things at once?
Not at small volumes. Every extra variant divides traffic and multiplies the visitors needed. On low traffic, one change per test, and one test at a time.
What should I test first?
Whatever everyone sees and sits close to the decision: the headline, the price, the wording of the offer. Testing an element seen by 5 percent of visitors takes twenty times the traffic for the same result.
Is a losing test wasted time?
No, it is preserved information: you know that avenue returns nothing, and you will not revisit it in six months. Record results, negative ones included, or you will re-run the same tests.