Statistical analysis
A two-proportion z-test on conversion, plus a Welch t-test on session duration. Nothing here is asserted beyond what the sample supports.
Variant B wins the split test
Variant B converted 17.0% better than Variant A. With a p-value of 0.0355, the sample provides clear evidence that this difference is real and not random noise.
Recommendation: Ship Variant B. At the observed 17.0% relative lift, rolling out the winning design should raise conversions on the same traffic volume.
- Winning variant
- Variant B
- Conversion difference
- +2.35% pts
- Relative lift
- +17.0%
- Confidence
- 96.5%
- p-value
- 0.0355
- Sample size
- 4,084
Confidence intervals
Wilson 95% intervals per arm — overlap means the difference is not yet resolved
Secondary test — session duration
Welch's t-test (unequal variances) on mean session duration
Mean session duration is 154.6s for Variant B versus 138.1s for Variant A — a difference of +16.5s (t = 0.99, p = 0.3215). This engagement gap is not statistically significant on the current sample.