Skip to content
Make Madly.
Tools / Testing
Calculator

A/B test significance calculator

Compare two conversion proportions with a p-value and interval, while seeing why sample size and uncertainty matter.

Last updated

Written by Madly editorial

01 / Your inputs

Run the numbers

Example figures are filled in so you can see a result straight away. Replace them with your own. Results update as you type; no data is sent or saved.

visitors

Visitors assigned to variant A.

conversions

Conversions from variant A; cannot exceed its visitors.

visitors

Visitors assigned to variant B.

conversions

Conversions from variant B; cannot exceed its visitors.

02 / Calculation

Observed difference (B − A)

Calculated result

0.8 percentage points

Unit: percentage points

Two-sided p-value
≈ 0.027303
95% CI lower (percentage points)
≈ 0.09 percentage points
95% CI upper (percentage points)
≈ 1.51 percentage points

Variant B minus variant A. The p-value uses a pooled two-proportion z-test; the 95% confidence interval uses an unpooled standard error. This is evidence, not a promise of a winner.

Statistical significance is not useful profit impact. Compare the extra conversions × contribution per conversion with incremental costs; this test does not calculate that business decision.

Formula

Rate = conversions ÷ visitors. For the two-sided p-value, z = (pB−pA) ÷ √[pPool(1−pPool)(1/nA+1/nB)]. For the approximate 95% interval, use (pB−pA) ± 1.96 × √[pA(1−pA)/nA + pB(1−pB)/nB].

How to read this

If the interval includes zero, the data remain compatible with either direction. Even if it excludes zero, check practical value, peeking, multiple tests and repeatability before changing spend.

Assumptions

Visitors are independent, mutually exclusive and randomly allocated. Use comparable attribution windows. For sparse expected counts the calculator withholds the p-value and interval because the normal approximation is unreliable; statistical evidence is not proof of a lasting effect.

How to use this tool

  • Enter visitors and conversions for version A.
  • Enter visitors and conversions for version B.
  • Read the difference in rate and the p-value.
  • Decide in advance what p-value you accept, and do not stop early because the number looks good.

Worked example

Illustrative example: 100 purchases from 2,000 visitors for A (5%) and 130 from 2,000 for B (6.5%) gives a 1.5 percentage-point observed difference; the interval and sample size matter more than the sign alone.

What this measures

Enter conversions and visitors for each version. The two-sided test uses a pooled proportion and the approximate interval uses separate proportions; both describe the observed difference B minus A.

Formula and units

Rate = conversions ÷ visitors. For the two-sided p-value, z = (pB−pA) ÷ √[pPool(1−pPool)(1/nA+1/nB)]. For the approximate 95% interval, use (pB−pA) ± 1.96 × √[pA(1−pA)/nA + pB(1−pB)/nB].

Assumptions and review notes

Visitors are independent, mutually exclusive and randomly allocated. Use comparable attribution windows. For sparse expected counts the calculator withholds the p-value and interval because the normal approximation is unreliable; statistical evidence is not proof of a lasting effect.

If the interval includes zero, the data remain compatible with either direction. Even if it excludes zero, check practical value, peeking, multiple tests and repeatability before changing spend.

A small p-value says the result would be unusual if there were no true difference. It does not say how big or how valuable the difference is. Many founders stop tests too early. If you check repeatedly and stop at the first good result, you will find false winners far more often than the p-value suggests.

  • Sample size.
  • Baseline conversion rate.
  • Size of the true difference.
  • How often you peek at results.
  • Stopping the moment p dips below 0.05.
  • Testing many variants and picking the best.
  • Ignoring practical size: a significant 0.1 point gain may not matter.
  • Using different traffic sources for A and B.

A closer look

Common questions.

What does significance mean?
It is a measure of how surprising the difference would be if both versions performed equally.
What p-value is good enough?
Many use 0.05, but choose a threshold before the test and consider the cost of being wrong.
Why is my result not significant?
Probably too few visitors or too small a difference. Run longer or test a bolder change.
Does significance mean it will keep winning?
No. It is a statement about this sample, and effects can shrink later.
How big a sample do I need?
Use the sample size calculator before the test to plan it.

Sources & further reading.

From reading to making

Make the next idea count.

Bring your product and your context. Explore a direction, then decide what belongs in the final ad.

Create account