4,084of 4,246 rows in scope
Statistics

Statistical analysis

A two-proportion z-test on conversion, plus a Welch t-test on session duration. Nothing here is asserted beyond what the sample supports.

Experiment verdictSignificant

Variant B wins the split test

Variant B converted 17.0% better than Variant A. With a p-value of 0.0355, the sample provides clear evidence that this difference is real and not random noise.

Recommendation: Ship Variant B. At the observed 17.0% relative lift, rolling out the winning design should raise conversions on the same traffic volume.

Winning variant
Variant B
Conversion difference
+2.35% pts
Relative lift
+17.0%
Confidence
96.5%
p-value
0.0355
Sample size
4,084
p-value
0.0355
Two-tailed, α = 0.05
z-score
2.103
Pooled standard error
Confidence
96.45%
1 − p
Observed power
55.6%
For the observed effect
Absolute difference
+2.35% pts
B − A
Relative lift
+16.95%
Difference ÷ baseline
95% CI of difference
+0.16% → +4.55%
Unpooled standard error
Required sample / arm
3,623
80% power at this effect size

Confidence intervals

Wilson 95% intervals per arm — overlap means the difference is not yet resolved

Variant A13.88% [12.46% , 15.43%]
Variant B16.23% [14.68% , 17.91%]

Secondary test — session duration

Welch's t-test (unequal variances) on mean session duration

Mean session duration is 154.6s for Variant B versus 138.1s for Variant A — a difference of +16.5s (t = 0.99, p = 0.3215). This engagement gap is not statistically significant on the current sample.