A/B testing compares two versions of a page, advert or user journey to find which performs better against a defined goal. This A B testing guide explains how to estimate sample sizes, interpret statistical significance and avoid mistakes that turn random variation into misleading marketing decisions.
A dependable test starts before either version goes live: choose one primary metric, define the smallest worthwhile improvement, calculate the traffic required and set a stopping rule. A higher conversion rate alone is not enough to declare a winner.
1. Write a test hypothesis and choose one primary metric
Begin with a specific friction point supported by analytics, user feedback or session observations. For example, visitors to an Islamabad consultancy’s enquiry page may abandon a form that asks for too much information before explaining what happens next.
Removing three optional fields will increase completed enquiries because visitors can request a consultation with less effort.
Keep the control unchanged and show the variant to a randomly selected audience. A 50/50 split typically gives the most statistical efficiency when traffic costs and variability are similar.
- Primary metric: Completed enquiries divided by eligible visitors.
- Guardrail: Qualified-lead rate, so additional spam or unsuitable enquiries do not count as commercial success.
- Secondary metric: Form starts, used to understand behaviour rather than select a winner.
- Randomisation unit: Usually the visitor, assigned consistently to one version where technically possible.
Define eligibility too. Exclude internal traffic and known bots using rules established before launch. Do not remove inconvenient results afterwards. If you test several changes together, you can assess the package, but not confidently attribute the outcome to one element.
2. Calculate sample size before launching
Sample size depends on your baseline conversion rate, minimum detectable effect, significance threshold and statistical power. Use a calculator designed for two independent conversion proportions, not a generic survey calculator.
- Estimate the baseline: Use recent, representative data from the same audience and funnel stage.
- Set the minimum detectable effect: Choose the smallest improvement worth implementing.
- Choose significance: A two-sided 5% threshold is a common starting point.
- Choose power: Typically 80% or 90%. Higher power requires more traffic.
- Calculate visitors per version: Do not mistake the per-variant requirement for the total.
Suppose your baseline is 3%, and you want to detect an increase to 3.6%. That is a 0.6 percentage-point absolute increase and a 20% relative increase. At 5% two-sided significance and 80% power, a standard approximation requires roughly 14,000 visitors per version, or 28,000 overall. Calculator methods can produce slightly different figures.
At 1,000 eligible visitors daily, reaching that total takes around four weeks. At 100 daily, it takes roughly nine months, during which seasonality and business changes may undermine the experiment’s relevance.
The practical lesson in this A B testing guide is that smaller effects need substantially more traffic. Halving the detectable absolute effect requires approximately four times the sample when other inputs remain similar. Do not choose an implausibly large uplift simply to make a calculator approve your traffic level.
3. Interpret significance, confidence intervals and business value
A p-value below 0.05 does not mean there is a 95% probability that your variant is better. It means that, under the no-difference hypothesis and the test’s assumptions, results at least as extreme as those observed would occur less than 5% of the time.
Report the estimated effect alongside its confidence interval. Imagine your observed uplift is 12%, but the interval ranges from a 3% decline to a 29% improvement. That result remains inconclusive: the data are compatible with both harm and benefit.
- Statistical significance: Is the evidence inconsistent with no difference under the chosen procedure?
- Practical significance: Is the likely benefit large enough to justify implementation?
- Uncertainty: Does the interval still include commercially unacceptable outcomes?
For a hypothetical Pakistan-based service business, assume an extra qualified lead contributes PKR 8,000 in expected gross profit after accounting for close rate. Ten additional qualified leads monthly would imply PKR 80,000 before implementation and ongoing costs. Use your own verified figures, not enquiry volume alone.
Agree the decision rule beforehand: ship only when the primary result meets the statistical criterion, the estimated benefit justifies costs and guardrails remain acceptable. A non-significant result is not proof that both versions perform identically.
4. Run the experiment through representative trading cycles
Traffic volume determines feasibility, but calendar coverage matters too. Typically, plan for at least two complete weekly cycles and the calculated sample, whichever takes longer. Two weeks is a planning floor, not a guarantee of validity.
For audiences in Islamabad or Lahore, Ramadan, Eid, salary dates and promotional campaigns may change buying behaviour. Running both variants concurrently helps balance these influences, but a seasonal result may not generalise to ordinary trading weeks.
Complete this launch checklist:
- Verify conversion events on mobile and desktop, including duplicate-event prevention.
- Check that assignment persists across repeat visits where possible.
- Confirm that neither variant introduces slower loading, broken forms or consent issues.
- Record campaign changes and avoid redesigning tested pages mid-experiment.
- Allow the same conversion window for both groups, especially when lead qualification takes several days.
If buying additional traffic, calculate costs first. For example, 28,000 incremental clicks at an assumed USD 0.50 per click would cost USD 14,000. That is an illustrative budget, not a market benchmark. Bought traffic may also behave differently from your normal organic audience.
5. Avoid the errors that create false winners
The most expensive testing mistakes often come from analysis choices rather than page design.
- Stopping at the first positive result: Repeatedly checking a fixed-horizon test and stopping when significant increases false positives. Follow the planned endpoint or use a properly configured sequential method.
- Ignoring unequal allocation: An unexplained departure from the expected traffic split can signal assignment or tracking problems. Investigate sample ratio mismatch before interpreting outcomes.
- Testing too many outcomes: Searching numerous metrics or audience segments for significance raises false-positive risk. Pre-specify comparisons and use appropriate multiple-testing corrections.
- Comparing different periods: Last month’s page versus this month’s redesign is not a randomised A/B test.
- Treating repeat visits as independent: Analyse consistently with the randomisation unit or use methods that account for dependence.
- Ignoring lead quality: More form submissions can still produce less revenue.
Keep an experiment log containing the hypothesis, dates, sample calculation, exclusions, results and rollout decision. This turns individual tests into reusable evidence rather than a collection of screenshots showing apparent winners.
Frequently asked questions
Can a low-traffic website run A/B tests?
Yes, but detecting modest improvements may take too long. Prioritise usability research, tracking repairs and stronger hypotheses first. Test a higher-volume funnel step only if it remains meaningfully connected to business outcomes.
Should every test use 95% significance?
No. This A B testing guide uses 95% as a common convention, not a universal rule. Choose thresholds before launch based on decision risk, and consider power and the cost of being wrong.
What should I do after an inconclusive result?
Review the confidence interval and implementation costs. Keep the control if evidence does not justify change. Do not extend a fixed-horizon test indefinitely; plan a fresh experiment or use a valid sequential approach.
Request a free SEO analysis from SEOISB, part of HA Technologies in Blue Area, Islamabad, to identify search visibility and conversion opportunities worth investigating.
