A/B Test Sample Size Calculator

Estimate the number of visitors needed per variant to detect a given lift with a chosen confidence and power.

The formula

Absolute lift δ = Baseline rate × Relative lift %
Treatment rate p₂ = p₁ + δ
n per variant = (z_α + z_β)² × (p₁(1 − p₁) + p₂(1 − p₂)) ÷ δ²
z_α from the two-sided confidence level, z_β from the power
Baseline rate p₁
Your current conversion rate as a percentage.
Relative lift
The smallest improvement worth detecting, expressed as a percentage of the baseline. A 15 percent lift on a 3.2 percent baseline is an absolute move to 3.68 percent.
Confidence
How rarely you are willing to call a difference real when there is none. 95 gives z of about 1.96.
Power
How likely you are to detect a real difference of the size you specified. 80 gives z of about 0.84.

Worked example

A 3.2 percent baseline, detecting a 15 percent relative lift, at 95 percent confidence and 80 percent power.

22,632 per variant, 45,264 in total, for a two-sided test with an even traffic split.

How to use it

  1. Enter your baseline conversion rate. What the current version already converts at, as a percentage. Take it from your own analytics over a full number of weeks, not a good day.
  2. Enter the lift you want to detect. A relative percentage. Entering 10 against a 5 percent baseline means you want to detect a move to 5.5 percent, not to 15 percent.
  3. Set confidence and power. Confidence is how often you would wrongly call a difference that is not there; power is how often you would find a real difference of that size. 95 and 80 are the usual starting points.
  4. Press Calculate Sample Size. You get the visitors needed per variant and the total across both. Divide the total by your weekly traffic to see how long the test has to run.

Visitors needed per variant, at 95% confidence and 80% power

Read down to your current conversion rate and across to the smallest lift worth detecting. Every figure was produced by this calculator’s own formula at those settings, and each is per variant — double it for the whole test.

Baseline+5% lift+10% lift+20% lift+50% lift
1%637,129163,12342,6997,749
2%315,26480,69421,1103,824
5%122,14431,2378,1571,469
10%57,77114,7513,839683
20%25,5856,5081,680291

The pattern to take away is that sample size grows roughly with the square of how small a lift you are chasing: detecting a 5 percent lift costs about four times the traffic of a 10 percent one. If the number for your baseline is out of reach, the honest options are to test a bolder change or to accept a lower confidence — not to stop the test early once it looks good.

Sample size scales with the inverse square of the effect

δ appears squared in the denominator, so halving the lift you want to detect roughly quadruples the traffic required. Detecting a 30 percent lift on the same baseline needs 6,035 per variant; detecting 5 percent needs 194,564. This is the single most important property of the calculation, and it is why tests aimed at small refinements on low-traffic pages never finish. If the required sample is out of reach, the productive response is to test a bolder change, not to run the same test for longer and hope.

Decide the sample before you start, then wait for it

Checking results daily and stopping when the difference looks significant inflates false positives far above the confidence level you selected, because you are effectively taking many looks at noisy data and stopping on the most flattering one. The number this calculator gives is a commitment: run to it, then decide. If you genuinely need to peek, that requires a sequential testing method, not an early stop on a fixed-horizon test.

Power is the half people skip

Confidence controls how often you see an effect that is not there; power controls how often you miss one that is. At 80 percent power, one real improvement in five goes undetected and gets recorded as no difference. Raising power to 90 percent costs roughly a third more traffic and is often worth it when the change is expensive to build or the decision is hard to revisit.

Time as well as traffic

Even when the sample arrives quickly, running for at least one full week — ideally two — keeps day-of-week effects from dominating. Weekday and weekend visitors behave differently, and a test that starts on a Monday and ends on a Thursday has sampled one kind of visitor. Divide the required total by your daily traffic to get the days needed, then round up to whole weeks.

Where to go next

Common questions

How many visitors do I need for an A/B test?

It depends on your baseline rate, the size of the effect you want to detect, and your chosen confidence and power. In the example — a 3.2 percent baseline, a 15 percent relative lift, 95 percent confidence, 80 percent power — it is about 22,600 per variant. Smaller effects need dramatically more.

Why does detecting a small lift need so much more traffic?

Because the effect size is squared in the denominator. Halving the lift you want to detect roughly quadruples the required sample. That relationship is why low-traffic pages can only realistically test large changes.

Can I stop the test as soon as it looks significant?

Not without breaking the statistics. Repeatedly checking and stopping at the first significant-looking moment produces false positives at a much higher rate than the confidence level suggests. Fix the sample size in advance and decide when you reach it.

What is the difference between confidence and power?

Confidence limits how often you declare a winner that is not real. Power is how likely you are to detect a real effect of the size you specified. At 80 percent power you miss one real improvement in five, which is why raising power matters when the change is costly to implement.

Get tools like this in your inbox
One useful tool per week. No spam. Unsubscribe anytime.
Tools Pro — $9 a month

One licence key, pasted into any CYZOR tool. Today it does three things: drops the CYZOR line from PDFs you send for signature, takes CYZOR branding off your forms and adds CSV export, and switches on the AI rewrite in the resume builder.

See Tools Pro

Terms · Privacy · Refund · Contact