Skip to content
Make Madly.
Tools / Testing
Calculator
New

A/B test sample size calculator

Plan visitors per test arm for a stated baseline, relative uplift, confidence and power.

Last updated

Calculated locally. Only these numeric inputs appear in a shared link.

Your results

Fixed-horizon, independent visitors and equal-allocation test assumptions
Target conversion rate
2.4%
Baseline multiplied by the relative uplift, not a percentage-point addition.
Visitors per arm
21,109
Equal allocation, two-sided fixed-horizon normal approximation.
Total eligible visitors
42,218
Control and treatment combined.
Estimated days
43
Traffic estimate only. Also cover normal business cycles.

How to use this tool

Use one reporting window and a consistent measurement scope. Replace the example inputs with your own figures. The result updates locally as you type. Copy the result to your planning notes, or copy a link containing only the numeric inputs. Reset restores the illustrative example.

  • Enter your current conversion rate.
  • Enter the smallest improvement worth detecting.
  • Choose the confidence and power levels.
  • Read the visitors needed per version and compare with your traffic.

Worked example

A 2% baseline and 20% relative uplift means detecting 2.4%, a 0.4 percentage-point change, not a rise to 22%. The calculation plans a two-sided, equal-allocation comparison of independent proportions at 95% confidence and 80% power.

Formula and units

p2 = p1 × (1 + relative uplift). p̄ = (p1 + p2) / 2. n per arm = [z(1-α/2)√(2p̄(1-p̄)) + z(power)√(p1(1-p1)+p2(1-p2))]² / (p2-p1)², rounded up.

Assumptions and review notes

Sample size is a planning approximation, not permission to stop when a dashboard turns green. Set the outcome and stopping rule before starting. Visitors should be independent and randomly assigned. Repeated views from the same person, multiple comparisons and platform delivery optimisation can break that assumption. Calendar days are a traffic estimate, not a guarantee.

A good plan decides sample size before the test starts and sticks to it. If you cannot reach the sample, test a bolder change.

  • Baseline rate.
  • Minimum detectable effect.
  • Confidence level.
  • Statistical power.
  • Stopping early.
  • Choosing an unrealistically small effect.
  • Ignoring traffic limits.
  • Running many variants.

A closer look

Common questions.

Why plan sample size?
To avoid ending tests before the data can tell versions apart.
What is minimum detectable effect?
The smallest change you care about.
What are power and confidence?
Chances of detecting a real effect and of avoiding a false alarm.
What if I lack traffic?
Test larger changes or use a higher level metric.
Does it store data?
No.

From reading to making

Make the next idea count.

Bring your product and your context. Explore a direction, then decide what belongs in the final ad.

Create account