A/B test significance calculator
Compare two conversion proportions with a p-value and interval, while seeing why sample size and uncertainty matter.
Last updated
Written by Madly editorial
01 / Your inputs
Run the numbers
Example figures are filled in so you can see a result straight away. Replace them with your own. Results update as you type; no data is sent or saved.
Visitors assigned to variant A.
Conversions from variant A; cannot exceed its visitors.
Visitors assigned to variant B.
Conversions from variant B; cannot exceed its visitors.
02 / Calculation
Observed difference (B − A)
Calculated result
0.8 percentage points
Unit: percentage points
- Two-sided p-value
- ≈ 0.027303
- 95% CI lower (percentage points)
- ≈ 0.09 percentage points
- 95% CI upper (percentage points)
- ≈ 1.51 percentage points
Variant B minus variant A. The p-value uses a pooled two-proportion z-test; the 95% confidence interval uses an unpooled standard error. This is evidence, not a promise of a winner.
Statistical significance is not useful profit impact. Compare the extra conversions × contribution per conversion with incremental costs; this test does not calculate that business decision.
Formula
Rate = conversions ÷ visitors. For the two-sided p-value, z = (pB−pA) ÷ √[pPool(1−pPool)(1/nA+1/nB)]. For the approximate 95% interval, use (pB−pA) ± 1.96 × √[pA(1−pA)/nA + pB(1−pB)/nB].
How to read this
If the interval includes zero, the data remain compatible with either direction. Even if it excludes zero, check practical value, peeking, multiple tests and repeatability before changing spend.
Assumptions
Visitors are independent, mutually exclusive and randomly allocated. Use comparable attribution windows. For sparse expected counts the calculator withholds the p-value and interval because the normal approximation is unreliable; statistical evidence is not proof of a lasting effect.
How to use this tool
- Enter visitors and conversions for version A.
- Enter visitors and conversions for version B.
- Read the difference in rate and the p-value.
- Decide in advance what p-value you accept, and do not stop early because the number looks good.
Worked example
Illustrative example: 100 purchases from 2,000 visitors for A (5%) and 130 from 2,000 for B (6.5%) gives a 1.5 percentage-point observed difference; the interval and sample size matter more than the sign alone.
What this measures
Enter conversions and visitors for each version. The two-sided test uses a pooled proportion and the approximate interval uses separate proportions; both describe the observed difference B minus A.
Formula and units
Rate = conversions ÷ visitors. For the two-sided p-value, z = (pB−pA) ÷ √[pPool(1−pPool)(1/nA+1/nB)]. For the approximate 95% interval, use (pB−pA) ± 1.96 × √[pA(1−pA)/nA + pB(1−pB)/nB].
Assumptions and review notes
Visitors are independent, mutually exclusive and randomly allocated. Use comparable attribution windows. For sparse expected counts the calculator withholds the p-value and interval because the normal approximation is unreliable; statistical evidence is not proof of a lasting effect.
If the interval includes zero, the data remain compatible with either direction. Even if it excludes zero, check practical value, peeking, multiple tests and repeatability before changing spend.
A small p-value says the result would be unusual if there were no true difference. It does not say how big or how valuable the difference is. Many founders stop tests too early. If you check repeatedly and stop at the first good result, you will find false winners far more often than the p-value suggests.
- Sample size.
- Baseline conversion rate.
- Size of the true difference.
- How often you peek at results.
- Stopping the moment p dips below 0.05.
- Testing many variants and picking the best.
- Ignoring practical size: a significant 0.1 point gain may not matter.
- Using different traffic sources for A and B.
A closer look
Common questions.
- What does significance mean?
- It is a measure of how surprising the difference would be if both versions performed equally.
- What p-value is good enough?
- Many use 0.05, but choose a threshold before the test and consider the cost of being wrong.
- Why is my result not significant?
- Probably too few visitors or too small a difference. Run longer or test a bolder change.
- Does significance mean it will keep winning?
- No. It is a statement about this sample, and effects can shrink later.
- How big a sample do I need?
- Use the sample size calculator before the test to plan it.