A/B testing (also called split testing) means showing two versions of something, a headline, an email subject line, a landing page, an ad creative, to different, randomly assigned groups of people, then comparing a metric between the groups to decide which version performs better. Done well, it replaces opinion with evidence. Done poorly, it produces confident-sounding conclusions that are actually noise.
The mechanics of running a test (a form field in your ad platform or a testing tool) are usually the easy part. The hard part is the thinking that should happen before and after: what exactly are you testing, how will you know it worked, and when are you allowed to stop and decide? This article covers that thinking.
Step 1: Write a specific hypothesis before you start
A hypothesis is a specific, falsifiable statement, not just 'let's test the headline.' A useful hypothesis names the change, the expected effect, and the reason you expect it. For example: 'Changing the landing page headline from a feature statement to an outcome statement will increase form submissions, because the current headline doesn't communicate the specific benefit visitors are searching for.' Writing this down before launch forces clarity on what you're actually testing and why, and it stops you from retroactively inventing a story for whatever the data shows.
Step 2: Choose one primary metric and guardrails
Every test needs exactly one primary metric that determines the decision, for example conversion rate, click-through rate, or cost per qualified lead. If you track ten metrics and declare victory whenever any one of them improves, you will find a 'winner' almost every time by chance alone, even with no real effect. Choose the primary metric that most directly reflects the business outcome you care about, then add one or two guardrail metrics to catch unintended harm, for example checking that a page redesign that improves clicks doesn't also tank the quality of the leads it generates.
Step 3: Randomize who sees each variant
A valid test requires random, comparable assignment, each visitor or recipient has a known chance of landing in variant A or B, and the split happens independently of anything about them (time of day, device, traffic source). Without randomization, apparent differences might reflect which audience saw which variant rather than the variant itself. Most ad platforms and testing tools randomize automatically; manual tests (like sending email A to one list segment and email B to another) need deliberate randomization, not a convenient split like 'first half of the list.'
Step 4: Understand sample size and why small tests mislead
Sample size is the number of visitors, sends, or impressions each variant needs before a result is trustworthy. Smaller samples produce noisier results: a 2-person test showing 1 conversion for A and 0 for B tells you nothing, even though A technically 'won.' Statistical significance calculators (many are free online, including ones built into major ad and email platforms) estimate how many conversions or visitors you need, given your current conversion rate and the minimum difference you'd consider meaningful. As a rule of thumb, tests on low-traffic pages or small email lists often need weeks, not days, to reach a reliable sample, which is why many small businesses are better served by directional learning (see below) than by formal statistical tests.
Step 5: Avoid peeking-driven conclusions
'Peeking' means checking results early and stopping the test the moment one variant looks ahead, rather than waiting for the pre-set sample size or time window. This inflates false positives because random fluctuations are larger early in a test and often reverse later. If you must look at results early for monitoring (for example, to catch a broken variant), separate 'checking for problems' from 'deciding a winner,' and commit to the pre-set stopping rule for the decision itself.
Directional learning versus statistically justified decisions
Not every business has the traffic to run a textbook significance test. It's honest to distinguish two tiers: a statistically justified decision, where you reached the planned sample size and a recognized method (like a chi-squared test or a platform's built-in significance calculation) confirms the result is unlikely to be chance; and directional learning, where a smaller test suggests a pattern worth exploring further but isn't proof. Both are useful, but they deserve different confidence and different-sized bets. Rolling out a major redesign site-wide based on a 40-visitor test is a bigger risk than piloting it further.
Statistical significance versus practical significance
A result can be statistically significant (unlikely to be due to chance) while being practically insignificant (too small to be worth the cost or complexity of change). A new checkout flow that lifts conversion from 3.00% to 3.05% might be statistically real on enough traffic, yet not worth the engineering time to maintain two flows. Before testing, it helps to set a minimum detectable effect, the smallest improvement that would actually be worth acting on, so a 'significant' tiny result doesn't automatically trigger a rollout.
A test-plan template you can reuse
- Hypothesis: the specific change, expected effect, and reasoning
- Primary metric: the single number that decides the test
- Guardrail metric(s): what must not get worse
- Minimum detectable effect: the smallest improvement worth acting on
- Sample size or duration: the pre-set stopping point, chosen before launch
- Randomization method: how variant assignment is made fair
- Decision rule: what counts as a win, a loss, or inconclusive
- Rollout plan: what happens next for each possible outcome
- 1Form a specific hypothesis and pick one primary metric
- 2Set sample size, duration, and decision rule before launch
- 3Randomize traffic and let the test run undisturbed
- 4Check guardrails for unintended harm, not to decide early
- 5Reach the pre-set stopping point, then evaluate and decide
- 6Document the result and the reasoning, win or lose
Common mistakes
- Testing too many elements at once (headline, image, and button color together) so you can't tell what caused the result.
- Stopping a test the moment it looks good, rather than at the pre-planned sample size.
- Declaring a 'winner' from a tiny sample with no significance check.
- Ignoring guardrail metrics and only watching the number you hope will go up.
- Running sequential tests back-to-back during different seasons or promotions and attributing the difference to the variant alone.
- Confusing statistical significance with a result that's actually big enough to matter for the business.
When this is not the right tactic
A/B testing is the wrong tool when traffic or sample size is too small to ever reach a meaningful result in a reasonable time, when the thing you're testing is a one-time decision rather than a repeatable page or message, or when you need to move fast on a low-stakes change that isn't worth the setup time. In those cases, qualitative research (customer interviews, usability observation), expert review, or simply making your best-informed decision and monitoring results over time is often more practical than forcing a formal test.
Your test-plan checklist
Before launching your next test, fill in the template above completely. If you cannot state a clear hypothesis, a single primary metric, and a stopping rule, you are not ready to launch the test yet, you are ready to think about it for another day.



