A/B Testing

A/B testing is a controlled experiment that randomly splits traffic between two versions of a product to measure which one changes a target metric.

15 min read

Β· Also in

By Ravi SuranaUpdated 5 sources

Quick answer

~20 sec

A/B testing is a controlled experiment that randomly assigns users to a current version (the control) or a changed version (the treatment) of a product, then compares a chosen metric between the two groups. Because assignment is random, any consistent difference between groups can be attributed to the change itself, not to who happened to see it.

011 min

A/B Testing at a glance

  • What it is: A randomized experiment that measures which version of something performs better on one clear metric.
  • Origin: Kohavi, Henne, and Sommerfield formalized web experimentation practice at Microsoft in a 2007 KDD paper.
  • Confused with: A multivariate test, which studies several changes and their interactions at once.
  • Watch for: Peeking at results early inflates the false-positive rate far past the intended 5%.

021 min

Why A/B Testing matters

A product manager watching a live test dashboard has to decide when a result is trustworthy enough to ship on. Get that call wrong and the company ships changes that never really worked, or kills changes that did work but had not run long enough to prove it. Watching a metric move after a launch cannot settle this by itself, because a holiday, a competitor's outage, or a marketing push landing the same week is an equally good explanation for the same move. A/B testing exists to answer the question directly, by comparing two groups that differ only in which version they saw.

Most of that confidence is misplaced to begin with. Only one third of the ideas tested on Microsoft's own experimentation platform improved the metric they were designed to improve, which is exactly why an objective measurement matters more than how good an idea sounds in a meeting. The cost of misreading it is concrete, not abstract. A growth team that ships on a peeked, not-yet-significant result can spend months building on a change that was actually neutral or harmful, and only find out once a later, properly run test contradicts it β€” after the engineering cost is already spent.

032 min

What A/B Testing actually measures

An A/B test splits traffic into two groups by chance: a control group that keeps the existing version, and a treatment group that sees the change. Every other factor β€” the day of the week, the mix of new and returning users, a marketing campaign running that week β€” lands on both groups in roughly equal proportion, because the split is random rather than chosen. That is what separates a controlled experiment from watching a metric move after a launch and guessing why. Correlation and Causation covers the general version of this gap: a metric changing after a release does not show the release caused it, because anything else that changed at the same time is an equally good explanation. A controlled experiment closes that gap by holding every other variable constant across the two groups through randomization alone, not through argument.

At the end of the planned run, the two groups' values on one chosen metric β€” a conversion rate, a click-through rate, a revenue figure β€” are compared with a statistical significance test, usually a two-proportion or two-mean test. The test produces a p-value: the probability of seeing a difference this large between the groups if the change actually had no effect at all. A result is typically called statistically significant when that p-value falls below a threshold fixed in advance, commonly 0.05. The threshold, the metric, and the number of users needed all have to be decided before the test starts. Deciding them afterward, once the data already point somewhere, is one of the ways a test stops being controlled.

041 min

Where A/B Testing comes from

Randomized experiments predate the web by decades β€” Ronald Fisher formalized the method for agricultural research in the 1920s β€” but running them continuously on live software traces to a specific paper. In 2007, Ron Kohavi, Randal Henne, and Dan Sommerfield of Microsoft published a paper at the ACM KDD conference making the case for measurement over opinion in software development. Kohavi, Henne, and Sommerfield argued that teams should listen to their customers, not to the Highest Paid Person's Opinion, or HiPPO. The term HiPPO stuck because it named something every team recognizes: a senior person's guess overriding a measured result. Their paper set out the statistical and engineering practices β€” randomization, sample size planning, variance reduction β€” that most A/B testing tools still implement today, and Kohavi went on to lead Microsoft's own Experimentation Platform, the source of much of the evidence in this article.

051 min

A worked example

Suppose a checkout page currently converts 4% of visitors into buyers, and the team wants to know whether a redesigned checkout button raises that rate by at least one percentage point, to 5%. Using the standard formula for a two-proportion test at 95% confidence and 80% power, that requires roughly 6,030 visitors in each group before the test can reliably tell a real one-point lift from noise.

If the team instead wants to detect a smaller, more realistic lift of half a percentage point β€” to 4.5% β€” the same formula requires roughly 24,110 visitors per group: about four times as many, because the required sample size grows with the square of how small an effect the test needs to catch.

Minimum detectable liftSample size needed per group
+1.0 percentage point (4% β†’ 5%)~6,030
+0.5 percentage point (4% β†’ 4.5%)~24,110

Halving the effect size does not double the required sample; it roughly quadruples it. A team that skips this calculation and just runs the test for a fixed two weeks, whatever traffic that happens to bring, often ends up with a result too noisy to trust in either direction.

061 min

How A/B Testing shows up in product and engineering work

A growth engineer ships a new onboarding flow behind a feature flag, splits new signups 50/50 between the old and new flow, and watches a dashboard update the split's activation rate hour by hour. Two days in, the new flow is ahead by three points, and someone on the team wants to roll it out to everyone immediately, before the planned two-week run finishes. This is the moment a controlled experiment either stays controlled or quietly stops being one: shipping on day two means shipping on a sample size the team never calculated as sufficient, using a result nobody planned to trust yet.

The same dashboard shows something else worth noticing: dozens of secondary metrics move along with the primary one, because a redesigned flow touches load time, error rate, and a handful of unrelated buttons on the same page. A PM who scans all of them looking for anything that moved is running dozens of informal tests inside one experiment, and a few of those will look significant by chance alone even when nothing real changed. The experimental mindset that makes A/B testing valuable β€” deciding the metric and the stopping point before looking at data β€” is also the discipline most under pressure the moment a dashboard starts moving in a direction somebody wanted.

071 min

When to use A/B Testing β€” and when not to

A/B testing fits a specific precondition: a metric that can be measured within the test's planned duration, a baseline value for that metric already known from existing traffic, and enough visitors passing through the experience to reach a sample size fixed before the test starts. Below a few thousand users per variant, most day-to-day product changes take weeks or months to reach a sample size large enough to detect anything but a huge effect, and the honest answer is that A/B testing is the wrong tool at that traffic level β€” a smaller-sample method such as a moderated usability study answers the question faster.

A/B testing also assumes one group's experience does not leak into the other's. A pricing change tested on one side of a marketplace can shift behavior on the other side even though that side saw no change at all, which breaks the assumption the test relies on. And it assumes the team already knows which metric decides success. Running the test first and deciding afterward which of the metrics that moved counts as the win is not A/B testing with an unclear goal β€” it is the multiple-comparisons problem, covered next, wearing a different name.

083 min

Where A/B Testing goes wrong

Four mechanisms account for most A/B tests that produce a confident, wrong answer.

Peeking at results before the planned sample size

Checking a test's p-value every day and stopping as soon as it crosses the significance threshold does not just risk an early call β€” it changes the actual false-positive rate. Evan Miller's analysis of this exact practice found that checking significance after every new observation and stopping the moment it looks significant pushes the effective false-positive rate to 26.1%, over five times the 5% a team thinks it agreed to. The pull to stop early is rarely just impatience β€” it is confirmation bias: a result trending toward what the team hoped for is far more tempting to lock in immediately than one trending the other way. The fix is deciding the sample size and the run length before the test starts, and treating any earlier look as informal.

The multiple comparisons problem

A test that tracks one primary metric and checks it once is asking one question. A test that tracks a hundred metrics, or gets re-run three times with small variations, is asking a hundred questions or three hundred, and at a 5% significance threshold, a few of those questions turn up "significant" purely by chance even when nothing real happened. Kohavi and Longbotham's own list of experiment pitfalls names this directly: teams need corrections for multiple hypothesis tests, because with enough treatments, metrics, or repeated experiment iterations, small p-values are likely to occur by chance alone.

Sample ratio mismatch

Before trusting any result, the experiment's own bookkeeping needs a check: did the number of users who actually landed in each group match the ratio the test was designed for? A test built for a 50/50 split that actually delivered 52/48 usually signals a bug in the assignment or logging, not a fluke. Fabijan and colleagues found that approximately 6% of experiments at Microsoft exhibit a sample ratio mismatch β€” common enough that mature experimentation platforms check for it automatically before showing anyone a result. A significant-looking lift sitting on top of a sample ratio mismatch is not evidence of anything; it is evidence the two groups were never comparable to begin with.

Statistically significant, practically meaningless

A test run on millions of users can find a real, statistically significant 0.02-percentage-point lift that would take a business a decade to notice in its revenue. Statistical significance says a difference is probably not noise. It says nothing about whether the difference is worth the engineering cost of shipping it, the added complexity of maintaining it, or the opportunity cost of not testing something bigger instead. Treating "significant" as a synonym for "ship it" mistakes a statistical property for a business judgment, and the two questions need to be asked separately.

091 min

A/B Testing vs. nearby concepts

The concept most often mistaken for A/B testing is the multivariate test. An A/B test compares two or more complete versions of an experience against each other. A multivariate test instead varies several elements at once β€” a headline, an image, a button color β€” independently, and measures not just each element's effect but how the elements interact with each other. The deciding fact is traffic: a multivariate test needs enough visitors to fill every combination of elements being varied, which multiplies the sample-size problem from the worked example above by the number of combinations, so it only works at traffic volumes far above what most A/B tests need.

A/B testing is also confused with an A/A test, which compares two identical versions against each other on purpose. An A/A test measures nothing about a real change β€” its only job is to confirm that the randomization and measurement system itself produces a result close to no difference when there is no real difference to find, which is exactly the kind of check that would have caught the Skype sample ratio mismatch below before it reached a real experiment.

102 min

A second case: what a trustworthy result looks like, and what a broken one looks like

The calculation above shows how to size a test properly. A real case shows what a trustworthy result looks like once the test is actually run. At Microsoft's Bing, a team added clickable links to individual sections of an advertiser's site directly inside the ad listing, then measured the change with a controlled experiment. The site links alone increased revenue but also degraded user metrics and page-load time; only after the team offset the extra space by showing fewer mainline ads per query did the feature end up improving revenue by tens of millions of dollars per year with neutral user impact. The full result needed both halves of that change to be true at once, which is exactly the kind of trade-off a single before/after number would have hidden.

A different case shows what happens when the measurement itself breaks. Fabijan and colleagues describe a Skype experiment that tested a machine-learning model for setting audio-buffering behavior against a fixed setting, with each call randomly assigned to one variant. In this Skype experiment, a logging bug meant the experiment analysis system collected 30% fewer sessions in one group than the design called for, triggering a sample ratio mismatch alert. The root cause traced back to the logging system: a variant ID tracked in memory kept updating mid-call even though the new setting itself was not applied mid-call, so a chunk of sessions ended up logged under the wrong variant. Every metric computed before that bug was found and fixed was comparing the wrong two groups.

This is why the discipline behind A/B testing includes checking the experiment's own bookkeeping, not just its headline result. A team that restricts its analysis to users who were still active weeks into the test runs into a related trap: analyzing users who have been active for a long time introduces survivorship bias, because the users who dropped out along the way are exactly the ones a fair comparison needs to keep.

112 min

Where the evidence on A/B Testing is contested

A controlled experiment is only as good as the metric it optimizes, and practitioners disagree about how hard that metric is to choose well. Kohavi and Longbotham warn that market share can be a long-term goal, but it is a terrible short-term criterion. A worse search engine can force people to issue more queries in the short run, which looks like more engagement even as users are quietly finding better alternatives elsewhere. The disagreement is not about whether controlled experiments work; it is about whether any short-run metric, measured over the one or two weeks a typical test runs, can be trusted to predict a change's effect months later.

Novelty and primacy effects sharpen the same worry. A visibly new design element often performs unusually well in its first days simply because it is new and draws extra attention β€” a primacy effect β€” before settling into whatever its steady-state performance actually is, and the reverse can happen when a familiar option is removed and users need time to adjust. A test that stops the moment an initial spike looks good, or the moment an initial dip looks bad, is measuring the novelty of the change rather than its real effect. Where practitioners land, in practice, is not abandoning short-run tests but running them past the point where the metric first looks decided, specifically to let a novelty spike fade before trusting the number underneath it.

122 min

Frequently asked questions about A/B Testing

What is A/B testing?

A/B testing is a controlled experiment that randomly splits users between a current version and a changed version of a product, then compares one chosen metric between the two groups to see whether the change caused a real difference.

What is the difference between A/B testing and multivariate testing?

A/B testing compares whole versions of an experience against each other. A multivariate test varies several elements independently and measures how they interact, which needs far more traffic to reach a reliable answer.

How do you calculate the sample size for an A/B test?

Using the baseline rate of the metric, the smallest lift worth detecting, and a target confidence and power β€” commonly 95% and 80% β€” a standard two-proportion formula gives the number of users needed per group before the test starts.

What is a good example of A/B testing?

Microsoft's Bing added clickable site links to ad listings and measured the change with a controlled experiment. The links alone increased revenue but also hurt user metrics and page-load time; only after the team offset the extra space with fewer ads per query did revenue rise by tens of millions of dollars a year with neutral user impact.

How do you avoid peeking bias in an A/B test?

Fix the sample size and run length before the test starts, and treat any result checked before that point as informal. Checking daily and stopping early can push the real false-positive rate to 26.1%, over five times the intended 5%.

What is a sample ratio mismatch?

A sample ratio mismatch is when the number of users actually assigned to each group does not match the ratio the test was designed for, which usually signals a bug in randomization or logging rather than a real result.

Why isn't a statistically significant result always worth shipping?

Statistical significance only says a difference is probably not random noise. It says nothing about whether that difference is large enough to justify the engineering and maintenance cost of shipping it.

?7 questions

Questions people ask

What is A/B testing?

A/B testing is a controlled experiment that randomly splits users between a current version and a changed version of a product, then compares one chosen metric between the two groups to see whether the change caused a real difference.

What is the difference between A/B testing and multivariate testing?

A/B testing compares whole versions of an experience against each other. A multivariate test varies several elements independently and measures how they interact, which needs far more traffic to reach a reliable answer.

How do you calculate the sample size for an A/B test?

Using the baseline rate of the metric, the smallest lift worth detecting, and a target confidence and power β€” commonly 95% and 80% β€” a standard two-proportion formula gives the number of users needed per group before the test starts.

What is a good example of A/B testing?

Microsoft's Bing added clickable site links to ad listings and measured the change with a controlled experiment. The links alone increased revenue but also hurt user metrics and page-load time; only after the team offset the extra space with fewer ads per query did revenue rise by tens of millions of dollars a year with neutral user impact.

How do you avoid peeking bias in an A/B test?

Fix the sample size and run length before the test starts, and treat any result checked before that point as informal. Checking daily and stopping early can push the real false-positive rate to 26.1%, over five times the intended 5%.

What is a sample ratio mismatch?

A sample ratio mismatch is when the number of users actually assigned to each group does not match the ratio the test was designed for, which usually signals a bug in randomization or logging rather than a real result.

Why isn't a statistically significant result always worth shipping?

Statistical significance only says a difference is probably not random noise. It says nothing about whether that difference is large enough to justify the engineering and maintenance cost of shipping it.

β–Ά3 videos

Watch

  • Understanding A/B Testing

  • Simple explanation of A/B Testing In Hindi

    codebasics Hindi

  • A/B Testing In Data Science

    Krish Naik

Β§5 sources

Sources

  1. Kohavi, R., & Longbotham, R. (2023). Online Controlled Experiments and A/B Tests. In Encyclopedia of Machine Learning and Data Science. Springer.

  2. Kohavi, R., Henne, R. M., & Sommerfield, D. (2007). Practical Guide to Controlled Experiments on the Web: Listen to Your Customers Not to the HiPPO. KDD 2007.

  3. Fabijan, A., Gupchup, J., Gupta, S., Omhover, J., Qin, W., Vermeer, L., & Dmitriev, P. (2019). Diagnosing Sample Ratio Mismatch in Online Controlled Experiments: A Taxonomy and Rules of Thumb for Practitioners. KDD 2019.

  4. Miller, E. How Not To Run an A/B Test. evanmiller.org.

Show all 5 sources

Keep reading

More from Research

All of Research
All of Research