Step 56 · AI Growth and Performance Foundations

A/B Testing Basics: Hypotheses, Sample Size, and Decision Rules

By the Daut Labz editorial teamPublished 7 min readintermediate

The short answer

A/B testing basics mean stating a hypothesis before you start, choosing one primary metric with guardrail metrics alongside it, randomizing who sees each variant, and waiting for a pre-set sample size or time window before deciding. The most common mistakes are stopping early when a result looks exciting ('peeking'), treating small differences as meaningful, and confusing statistical significance with practical significance that's actually worth acting on.

Two hand-drawn creative variants, A and B, each pointing into a shared funnel labeled with a single measurement plan.

Key takeaways

  • Write your hypothesis and primary metric before launching the test, not after you see results.
  • A result can be statistically significant but not practically significant enough to matter.
  • Checking results daily and stopping as soon as one variant 'wins' inflates false positives.
  • Guardrail metrics catch cases where a test improves one number by quietly harming another.
  • Underpowered tests (too few visitors or conversions) often produce noise that looks like a signal.

Helpful first: Test Meta Ad Creative With a Clear Hypothesis

A/B testing (also called split testing) means showing two versions of something, a headline, an email subject line, a landing page, an ad creative, to different, randomly assigned groups of people, then comparing a metric between the groups to decide which version performs better. Done well, it replaces opinion with evidence. Done poorly, it produces confident-sounding conclusions that are actually noise.

The mechanics of running a test (a form field in your ad platform or a testing tool) are usually the easy part. The hard part is the thinking that should happen before and after: what exactly are you testing, how will you know it worked, and when are you allowed to stop and decide? This article covers that thinking.

Step 1: Write a specific hypothesis before you start

A hypothesis is a specific, falsifiable statement, not just 'let's test the headline.' A useful hypothesis names the change, the expected effect, and the reason you expect it. For example: 'Changing the landing page headline from a feature statement to an outcome statement will increase form submissions, because the current headline doesn't communicate the specific benefit visitors are searching for.' Writing this down before launch forces clarity on what you're actually testing and why, and it stops you from retroactively inventing a story for whatever the data shows.

Step 2: Choose one primary metric and guardrails

Every test needs exactly one primary metric that determines the decision, for example conversion rate, click-through rate, or cost per qualified lead. If you track ten metrics and declare victory whenever any one of them improves, you will find a 'winner' almost every time by chance alone, even with no real effect. Choose the primary metric that most directly reflects the business outcome you care about, then add one or two guardrail metrics to catch unintended harm, for example checking that a page redesign that improves clicks doesn't also tank the quality of the leads it generates.

Step 3: Randomize who sees each variant

A valid test requires random, comparable assignment, each visitor or recipient has a known chance of landing in variant A or B, and the split happens independently of anything about them (time of day, device, traffic source). Without randomization, apparent differences might reflect which audience saw which variant rather than the variant itself. Most ad platforms and testing tools randomize automatically; manual tests (like sending email A to one list segment and email B to another) need deliberate randomization, not a convenient split like 'first half of the list.'

Step 4: Understand sample size and why small tests mislead

Sample size is the number of visitors, sends, or impressions each variant needs before a result is trustworthy. Smaller samples produce noisier results: a 2-person test showing 1 conversion for A and 0 for B tells you nothing, even though A technically 'won.' Statistical significance calculators (many are free online, including ones built into major ad and email platforms) estimate how many conversions or visitors you need, given your current conversion rate and the minimum difference you'd consider meaningful. As a rule of thumb, tests on low-traffic pages or small email lists often need weeks, not days, to reach a reliable sample, which is why many small businesses are better served by directional learning (see below) than by formal statistical tests.

Step 5: Avoid peeking-driven conclusions

'Peeking' means checking results early and stopping the test the moment one variant looks ahead, rather than waiting for the pre-set sample size or time window. This inflates false positives because random fluctuations are larger early in a test and often reverse later. If you must look at results early for monitoring (for example, to catch a broken variant), separate 'checking for problems' from 'deciding a winner,' and commit to the pre-set stopping rule for the decision itself.

Directional learning versus statistically justified decisions

Not every business has the traffic to run a textbook significance test. It's honest to distinguish two tiers: a statistically justified decision, where you reached the planned sample size and a recognized method (like a chi-squared test or a platform's built-in significance calculation) confirms the result is unlikely to be chance; and directional learning, where a smaller test suggests a pattern worth exploring further but isn't proof. Both are useful, but they deserve different confidence and different-sized bets. Rolling out a major redesign site-wide based on a 40-visitor test is a bigger risk than piloting it further.

Statistical significance versus practical significance

A result can be statistically significant (unlikely to be due to chance) while being practically insignificant (too small to be worth the cost or complexity of change). A new checkout flow that lifts conversion from 3.00% to 3.05% might be statistically real on enough traffic, yet not worth the engineering time to maintain two flows. Before testing, it helps to set a minimum detectable effect, the smallest improvement that would actually be worth acting on, so a 'significant' tiny result doesn't automatically trigger a rollout.

A test-plan template you can reuse

A/B test plan template
  • Hypothesis: the specific change, expected effect, and reasoning
  • Primary metric: the single number that decides the test
  • Guardrail metric(s): what must not get worse
  • Minimum detectable effect: the smallest improvement worth acting on
  • Sample size or duration: the pre-set stopping point, chosen before launch
  • Randomization method: how variant assignment is made fair
  • Decision rule: what counts as a win, a loss, or inconclusive
  • Rollout plan: what happens next for each possible outcome
The A/B testing decision cycle
  1. 1Form a specific hypothesis and pick one primary metric
  2. 2Set sample size, duration, and decision rule before launch
  3. 3Randomize traffic and let the test run undisturbed
  4. 4Check guardrails for unintended harm, not to decide early
  5. 5Reach the pre-set stopping point, then evaluate and decide
  6. 6Document the result and the reasoning, win or lose

Common mistakes

  • Testing too many elements at once (headline, image, and button color together) so you can't tell what caused the result.
  • Stopping a test the moment it looks good, rather than at the pre-planned sample size.
  • Declaring a 'winner' from a tiny sample with no significance check.
  • Ignoring guardrail metrics and only watching the number you hope will go up.
  • Running sequential tests back-to-back during different seasons or promotions and attributing the difference to the variant alone.
  • Confusing statistical significance with a result that's actually big enough to matter for the business.

When this is not the right tactic

A/B testing is the wrong tool when traffic or sample size is too small to ever reach a meaningful result in a reasonable time, when the thing you're testing is a one-time decision rather than a repeatable page or message, or when you need to move fast on a low-stakes change that isn't worth the setup time. In those cases, qualitative research (customer interviews, usability observation), expert review, or simply making your best-informed decision and monitoring results over time is often more practical than forcing a formal test.

Your test-plan checklist

Before launching your next test, fill in the template above completely. If you cannot state a clear hypothesis, a single primary metric, and a stopping rule, you are not ready to launch the test yet, you are ready to think about it for another day.

Frequently asked questions

What is a good primary metric for an A/B test?

Choose the metric closest to the business outcome you care about, for example conversion rate or qualified-lead rate, rather than a vanity metric like page views. Pick only one primary metric per test so the decision rule stays unambiguous.

How long should an A/B test run?

Long enough to reach your pre-set sample size, often at least one to two full business cycles (commonly one to two weeks) to average out day-of-week effects. There is no universal fixed number of days that works for every business.

What does 'peeking' mean in A/B testing?

Peeking means checking results early and stopping as soon as one variant appears to win, rather than waiting for the planned sample size. It inflates the chance of a false positive because early results are noisier than they look.

Is a statistically significant result always worth acting on?

Not necessarily. A result can be statistically significant but too small in practical terms to justify the cost of changing anything. Set a minimum detectable effect before testing so you know what size of result would actually be worth acting on.

What if I don't have enough traffic for a formal significance test?

Treat smaller tests as directional learning rather than statistically justified decisions. Use them to inform smaller, lower-risk bets, and consider qualitative methods like customer interviews alongside or instead of formal testing.

Sources

Related guides

A hand-drawn storyboard of several ad creative variants arranged beside an open experiment notebook with notes.

Meta Ads and Paid Media

Step 46

Test Meta Ad Creative With a Clear Hypothesis

A structured approach to testing Meta ad creative: separate concept, hook, proof, format, and offer, test one meaningful change at a time, and judge results against a predefined decision metric, not a feeling.

  • Meta ads
  • creative testing
  • ad experiments
  • hypothesis testing
7 min readintermediate
Read →
Two comparable hand-drawn groups of customers separated by a clear ink line representing an experimental boundary.

Advanced Growth and Measurement

Step 71

Incrementality Testing: Did Marketing Cause Additional Business?

A pro-level introduction to incrementality testing: why attributed conversions are not proof of causality, how randomized and geo holdouts work, and a worked hypothetical lift calculation you can adapt.

  • incrementality
  • holdout testing
  • measurement
  • pro
6 min readpro
Read →
A hand-drawn landing-page testing lab with controlled page variants side by side and an analyst reviewing results in a notebook.

Advanced Growth and Measurement

Step 77

Advanced Conversion Experiments: Prioritization, Power, and Guardrails

A pro-level framework for running credible conversion experiments: diagnosing real friction, prioritizing hypotheses, sizing tests correctly, and turning inconclusive results into useful learning instead of wasted traffic.

  • cro
  • experimentation
  • a/b testing
  • statistics
7 min readpro
Read →