performancepositive
Variant B is currently ahead
16.23%
Variant B conversion rate
Variant B converts at 16.23% versus 13.88% for Variant A — a relative lift of 17.0% on the current sample.
Why it matters — The difference clears the 95% significance threshold, so it is unlikely to be noise.
Explore Variant B →statisticspositive
Result is statistically significant (p = 0.0355)
96.5%
Confidence level
Variant B converted 17.0% better than Variant A. With a p-value of 0.0355, the sample provides clear evidence that this difference is real and not random noise.
Why it matters — You can act on the winning variant with a conventional 5% false-positive risk.
Open statistical analysis →behaviorneutral
Longer sessions convert substantially better
0.042
Correlation with conversion
Sessions in the 600s+ band convert at 31.43% compared with 12.72% in the 0–30s band (point-biserial r = 0.042).
Why it matters — Engagement depth is a leading indicator: if the winning design also raises dwell time, the lift is likely behavioural rather than random.
Open behaviour analysis →behaviorneutral
Variant B holds users 17s longer per session
+17s
Mean duration difference
Mean session duration is 155s for Variant B versus 138s for Variant A. Welch's t-test returns p = 0.3215.
Why it matters — A durable engagement gap supports the conversion result and suggests the design change alters behaviour, not just outcomes.
Compare distributions →segmentneutral
tablet traffic converts best, mobile lags
3.73%
Segment gap
tablet sessions convert at 17.07% (334 sessions) while mobile converts at 13.34%.
Why it matters — Device mix imbalance between arms can distort the headline result — check that traffic split is even before shipping.
Inspect segments →segmentneutral
organic is the largest acquisition channel
1,210
Sessions
organic accounts for 29.6% of all sessions in the experiment and converts at 13.80%.
Why it matters — The dominant channel drives most of the headline number; a shift in its mix can move the experiment result on its own.
Open data explorer →qualitywarning
236 session-duration outliers detected
236
Outlier rows
Using the 1.5×IQR rule, 5.56% of rows sit far outside the typical duration range — likely abandoned tabs or bot traffic.
Why it matters — Extreme durations inflate mean-based metrics. Median duration and the outlier toggle in the filter bar guard against this.
Review data quality →qualitywarning
Dataset needs sanitisation before reporting
91/100
Data quality score
46 duplicate records and 199 missing cells were found across 4,246 rows (99.61% complete).
Why it matters — Duplicates double-count conversions and bias the split. All dashboard metrics exclude them by default.
Open data quality →