A/B Test Sample Size Calculator
Estimate the number of visitors needed per variant to detect a given lift with a chosen confidence and power.
Estimate the number of visitors needed per variant to detect a given lift with a chosen confidence and power.
A 3.2 percent baseline, detecting a 15 percent relative lift, at 95 percent confidence and 80 percent power.
22,632 per variant, 45,264 in total, for a two-sided test with an even traffic split.
Read down to your current conversion rate and across to the smallest lift worth detecting. Every figure was produced by this calculator’s own formula at those settings, and each is per variant — double it for the whole test.
| Baseline | +5% lift | +10% lift | +20% lift | +50% lift |
|---|---|---|---|---|
| 1% | 637,129 | 163,123 | 42,699 | 7,749 |
| 2% | 315,264 | 80,694 | 21,110 | 3,824 |
| 5% | 122,144 | 31,237 | 8,157 | 1,469 |
| 10% | 57,771 | 14,751 | 3,839 | 683 |
| 20% | 25,585 | 6,508 | 1,680 | 291 |
The pattern to take away is that sample size grows roughly with the square of how small a lift you are chasing: detecting a 5 percent lift costs about four times the traffic of a 10 percent one. If the number for your baseline is out of reach, the honest options are to test a bolder change or to accept a lower confidence — not to stop the test early once it looks good.
δ appears squared in the denominator, so halving the lift you want to detect roughly quadruples the traffic required. Detecting a 30 percent lift on the same baseline needs 6,035 per variant; detecting 5 percent needs 194,564. This is the single most important property of the calculation, and it is why tests aimed at small refinements on low-traffic pages never finish. If the required sample is out of reach, the productive response is to test a bolder change, not to run the same test for longer and hope.
Checking results daily and stopping when the difference looks significant inflates false positives far above the confidence level you selected, because you are effectively taking many looks at noisy data and stopping on the most flattering one. The number this calculator gives is a commitment: run to it, then decide. If you genuinely need to peek, that requires a sequential testing method, not an early stop on a fixed-horizon test.
Confidence controls how often you see an effect that is not there; power controls how often you miss one that is. At 80 percent power, one real improvement in five goes undetected and gets recorded as no difference. Raising power to 90 percent costs roughly a third more traffic and is often worth it when the change is expensive to build or the decision is hard to revisit.
Even when the sample arrives quickly, running for at least one full week — ideally two — keeps day-of-week effects from dominating. Weekday and weekend visitors behave differently, and a test that starts on a Monday and ends on a Thursday has sampled one kind of visitor. Divide the required total by your daily traffic to get the days needed, then round up to whole weeks.
It depends on your baseline rate, the size of the effect you want to detect, and your chosen confidence and power. In the example — a 3.2 percent baseline, a 15 percent relative lift, 95 percent confidence, 80 percent power — it is about 22,600 per variant. Smaller effects need dramatically more.
Because the effect size is squared in the denominator. Halving the lift you want to detect roughly quadruples the required sample. That relationship is why low-traffic pages can only realistically test large changes.
Not without breaking the statistics. Repeatedly checking and stopping at the first significant-looking moment produces false positives at a much higher rate than the confidence level suggests. Fix the sample size in advance and decide when you reach it.
Confidence limits how often you declare a winner that is not real. Power is how likely you are to detect a real effect of the size you specified. At 80 percent power you miss one real improvement in five, which is why raising power matters when the change is costly to implement.
One licence key, pasted into any CYZOR tool. Today it does three things: drops the CYZOR line from PDFs you send for signature, takes CYZOR branding off your forms and adds CSV export, and switches on the AI rewrite in the resume builder.