A/B test sample size calculator
Plan visitors per test arm for a stated baseline, relative uplift, confidence and power.
Last updated
Your results
Fixed-horizon, independent visitors and equal-allocation test assumptions- Target conversion rate
- 2.4%
- Baseline multiplied by the relative uplift, not a percentage-point addition.
- Visitors per arm
- 21,109
- Equal allocation, two-sided fixed-horizon normal approximation.
- Total eligible visitors
- 42,218
- Control and treatment combined.
- Estimated days
- 43
- Traffic estimate only. Also cover normal business cycles.
How to use this tool
Use one reporting window and a consistent measurement scope. Replace the example inputs with your own figures. The result updates locally as you type. Copy the result to your planning notes, or copy a link containing only the numeric inputs. Reset restores the illustrative example.
- Enter your current conversion rate.
- Enter the smallest improvement worth detecting.
- Choose the confidence and power levels.
- Read the visitors needed per version and compare with your traffic.
Worked example
A 2% baseline and 20% relative uplift means detecting 2.4%, a 0.4 percentage-point change, not a rise to 22%. The calculation plans a two-sided, equal-allocation comparison of independent proportions at 95% confidence and 80% power.
Formula and units
p2 = p1 × (1 + relative uplift). p̄ = (p1 + p2) / 2. n per arm = [z(1-α/2)√(2p̄(1-p̄)) + z(power)√(p1(1-p1)+p2(1-p2))]² / (p2-p1)², rounded up.
Assumptions and review notes
Sample size is a planning approximation, not permission to stop when a dashboard turns green. Set the outcome and stopping rule before starting. Visitors should be independent and randomly assigned. Repeated views from the same person, multiple comparisons and platform delivery optimisation can break that assumption. Calendar days are a traffic estimate, not a guarantee.
A good plan decides sample size before the test starts and sticks to it. If you cannot reach the sample, test a bolder change.
- Baseline rate.
- Minimum detectable effect.
- Confidence level.
- Statistical power.
- Stopping early.
- Choosing an unrealistically small effect.
- Ignoring traffic limits.
- Running many variants.
A closer look
Common questions.
- Why plan sample size?
- To avoid ending tests before the data can tell versions apart.
- What is minimum detectable effect?
- The smallest change you care about.
- What are power and confidence?
- Chances of detecting a real effect and of avoiding a false alarm.
- What if I lack traffic?
- Test larger changes or use a higher level metric.
- Does it store data?
- No.