Learn how to define a testable hypothesis, calculate sample size, set the runtime and interpret results before you trust an A/B test.
Use an A/B test when you have a specific change, a measurable target action and enough traffic to reach the required sample within a useful period. It works well for questions such as whether a revised checkout step improves completion. It is the wrong method when the question is vague, several major elements change at once or traffic is too low for a conclusive result.
A useful hypothesis states the observed problem, its likely cause, the proposed change and the expected effect on one primary metric. For example: “Visitors overlook delivery details below the fold. Moving them next to the price will increase the add-to-cart rate.” Base the observation on behavioural or funnel data rather than personal preference.
Prioritise reliable random assignment, mutually exclusive variants, consistent user allocation and clear reporting of sample size, confidence intervals and effect size. Check whether the tool works with your analytics setup and consent process under GDPR. Visual editors and templates matter less if you cannot verify exposure, data quality or how the result was calculated.
Set the sample size before launch using your baseline conversion rate, minimum detectable effect, significance level and desired statistical power. Smaller effects require larger samples. Calculate the requirement for each variant, not just the test as a whole, and confirm that the relevant page or funnel receives enough eligible traffic to reach it.
Run the test until the planned sample is reached and include complete business cycles, such as weekdays and weekends. Avoid choosing an end date only because the dashboard shows a winner. Also account for campaign schedules, holidays and unusual outages. If reaching the sample would take so long that conditions are likely to change, testing may not be practical.
Statistical significance estimates how compatible the observed difference is with there being no real difference between the variants, given the model and assumptions. It does not show whether the improvement is large enough to matter. Review the effect size and confidence interval too, then weigh the likely gain against implementation cost, risk and impact on secondary metrics.
Do not lower the evidence threshold or leave an underpowered test running indefinitely. Use Smart Funnels to locate drop-off, Heatmaps to assess attention and interaction, and Session Replays to inspect friction such as Rage Clicks. You can also run user interviews or usability tests. These methods do not prove uplift, but they can reveal stronger problems and support a better-informed change.