What is A/B testing?
A/B testing shows two versions of a page to random halves of your visitors and measures which one earns more. How it works, and when it is worth doing.
Updated 28 September 2026 · 7 min read
A/B testing splits visitors at random between the current version of a page (A) and a changed version (B), then compares one result that matters, such as orders or revenue per visitor. Because chance decides who sees what, a big enough gap can be pinned on the change. At a 3% conversion rate you need about 28,000 visitors in total to reliably spot a 20% lift, so low-traffic sites should test bold changes.
A/B testing is a controlled experiment on your own visitors. You keep the current version of something (a landing page, a pricing page, a checkout step, an email subject line) as version A, build one change as version B, and let a coin flip decide which version each visitor sees. After enough visitors, you compare the two groups on one number you care about, like orders or revenue per visitor.
The coin flip is what makes it work. Random assignment means the two groups match on everything except your change: the same mix of phones and laptops, the same ad campaigns, the same days of the week. So if B's group buys more, and the gap is bigger than chance alone would produce, the change caused it. No analytics dashboard can tell you that on its own.
How an A/B test works
- Pick the page and the single change you want to test.
- Decide which metric picks the winner, before the test starts. For a store or a SaaS product, revenue per visitor is usually the right one (see /guides/revenue-per-visitor).
- Work out how many visitors each version needs to detect the smallest lift you care about. The /tools/ab-test-sample-size-calculator does this in a few seconds.
- Split visitors at random, usually 50/50, and keep each person in the same group on every visit.
- Run the test to the planned sample size, in whole weeks, without stopping when it looks good.
- Compare the groups once, ship the winner, or keep the original if there is no clear difference.
The full process, with checks at each stage, is in /guides/how-to-run-an-ab-test.
A worked example
Say an online store's product page turns 3.0% of visitors into buyers. The team thinks showing a delivery date next to the "Add to cart" button will help. They split 20,000 visitors evenly. These are example numbers:
| Version | Visitors | Orders | Conversion rate |
|---|---|---|---|
| A (current page) | 10,000 | 300 | 3.0% |
| B (delivery date shown) | 10,000 | 360 | 3.6% |
The lift is (3.6 - 3.0) / 3.0 = 20%. Is that real, or luck? A two-proportion z-test answers it:
- Pooled conversion rate: 660 / 20,000 = 3.3%
- Standard error of the difference: √(0.033 × 0.967 × (1/10,000 + 1/10,000)) = 0.00253
- z = (0.036 - 0.030) / 0.00253 = 2.38
- Two-sided p-value = 0.018
A p-value of 0.018 means that if the two pages truly performed the same, a gap this large would show up about 1.8% of the time. That is below the usual 0.05 bar, so the team ships B. The 95% confidence interval for the lift runs from about +3.5% to +36.5%, which tells you B is very likely better but the true gain could be far smaller than the 20% measured. You can run the same numbers in the /tools/ab-test-significance-calculator, and /guides/statistical-significance-explained covers what the p-value does and doesn't mean.
Why not just compare before and after?
Launching the change and comparing this month with last month feels simpler. The problem is that many things move a conversion rate at once: seasonality, a new ad campaign, a competitor's sale, a search ranking change, payday. A before-and-after comparison can't separate your change from any of them. In an A/B test both groups live through the same days, so outside events hit both equally and cancel out.
Testing also catches ideas that look minor but aren't. In 2012 a Bing employee proposed a small change to how ad headlines were displayed. It sat in the backlog for more than six months because it seemed low priority. When an engineer finally ran it as an A/B test, it raised revenue by 12%, worth more than $100 million a year in the US alone, according to Kohavi and Thomke in Harvard Business Review.
Most ideas don't win
The same HBR article reports that only about 10 to 20% of experiments at Google and Bing generate positive results. Across Microsoft, roughly one third of experiments improve the metric they were designed to improve, one third have no effect and one third make it worse. An earlier Microsoft paper, Online Experimentation at Microsoft, found the same one-third success rate for well-designed experiments.
This is the strongest argument for testing. If experienced product teams at Microsoft are wrong about two thirds of their ideas, a founder's gut feel about a new pricing page is not a safe bet either. Testing turns "we think" into "we measured", and it stops you from shipping the two thirds of changes that do nothing or hurt.
It is also a warning. Most tests will end with no winner, and a program that "wins" every test is usually calling noise a win.
When A/B testing is worth it
A/B testing is worth running when three things are true:
- The page gets enough traffic to detect a lift you'd care about within a few weeks.
- The decision is genuinely uncertain. If a page is broken, fix it; don't test it.
- The metric you care about shows up within the test window. Purchases and signups do. Twelve-month retention does not.
Traffic is usually the deciding factor. This table shows the smallest relative lift you can reliably detect (95% confidence, 80% power) on a 3% conversion rate if you run for four weeks and split traffic 50/50:
| Visitors a day | Visitors per version after 4 weeks | Smallest detectable lift |
|---|---|---|
| 500 | 7,000 | about 29% |
| 1,000 | 14,000 | about 20% |
| 2,000 | 28,000 | about 14% |
| 4,000 | 56,000 | about 10% |
| 10,000 | 140,000 | about 6% |
A site with 500 visitors a day can still test, but only changes that could plausibly lift conversion by 30% or more: a new offer, a different price, a shorter checkout, a rewritten headline that changes what you promise. Button colors will never reach significance. /guides/ab-testing-low-traffic covers the options when traffic is thin.
In A/B Testing Intuition Busters, Ronny Kohavi, Alex Deng and Lukas Vermeer give the general guidance that A/B tests are useful for effects of reasonable size when you have at least thousands of active users, and preferably tens of thousands.
When it isn't worth it
Skip the test, or pick a different method, when:
- You have a known bug, a broken form or a page that fails on mobile. Fix it.
- The change is required anyway, like a legal notice or a price rise your costs force on you.
- The page gets a few hundred visitors a month. You'll wait a year for an answer. Talk to customers, watch session recordings and make the call.
- The effect takes months to show, such as long-term churn, and you can't measure a leading indicator in the meantime.
- You want to test many small things at once on modest traffic. Each extra version splits your traffic further (see /guides/ab-testing-vs-multivariate-testing).
What makes a result trustworthy
A test can be run perfectly and still mislead you if the setup is wrong. The common failures:
- Stopping the test the first day it looks significant. Checking daily and stopping at the first significant result can push the false win rate from 5% to over 25% (see /guides/peeking-problem-ab-testing).
- Running without a planned sample size, so the test is too small to detect anything real.
- Judging on clicks or signups when the business runs on revenue. A version can win signups and lose money.
- A broken traffic split. If you asked for 50/50 and got 52/48 on 20,000 visitors, something is filtering visitors out of one version. /guides/sample-ratio-mismatch explains how to check.
- Stopping mid-week. Weekday and weekend visitors often behave differently, so run whole weeks.
/guides/ab-testing-mistakes goes through each of these with fixes.
Terms you'll see
| Term | What it means |
|---|---|
| A/B test | One changed version against the original |
| A/B/n test | Several changed versions against the original, each getting a share of traffic |
| Multivariate test | Several elements changed at once in every combination, to measure each element's effect |
| A/A test | The original against itself, to check that the testing setup doesn't find differences that aren't there |
| Split URL test | An A/B test where the versions live at different web addresses |
| Holdout | A small group kept on the original after a winner ships, to measure what the winner really earned |
Doing it without a data team
The method is simple. The discipline around it (picking the right metric, sizing the test, not peeking, checking the split, keeping a holdout) is where most teams slip. Outtest runs that loop. It connects read-only to your analytics and payment tools, finds where the funnel loses the most money, launches the test, and calls a winner only when it clears three bars on revenue per visitor. You can also do all of it by hand with the free calculators linked above. The maths is the same either way.
Questions people ask
What is A/B testing in simple terms?+
You show the current version of a page to half your visitors and a changed version to the other half, picked at random. After enough visitors, you compare which group bought more (or signed up more, or earned more per visitor). Random assignment means the change is the only systematic difference between the groups, so a large enough gap was caused by the change.
Is A/B testing the same as split testing?+
Yes, the terms are used for the same thing. Some tools use 'split URL testing' for the case where A and B live at different web addresses, but the method and the statistics are the same.
How much traffic do you need for A/B testing?+
It depends on your conversion rate and the size of lift you want to detect. At a 3% conversion rate, detecting a 20% relative lift at 95% confidence and 80% power takes about 13,900 visitors per version, or 27,800 in total. A site with 1,000 visitors a day gets there in four weeks; a site with 200 a day would need over four months.
How often do A/B tests find a winner?+
Less often than people expect. Ronny Kohavi and Stefan Thomke reported in Harvard Business Review that only about 10 to 20% of experiments at Google and Bing generate positive results, and about one third at Microsoft overall. Plan for most tests to end with no clear winner.
Is A/B testing worth it for a small business?+
It is worth it once a page gets a few thousand visitors a month and you test changes big enough to move the number by 20% or more. Below that, you will wait months for an answer, so fix obvious problems directly and save testing for decisions you are genuinely unsure about.
Read next
Let Outtest run your split tests
AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.