Split testing guides
How to run A/B tests that pay: what to test, how many visitors you need, when to call a winner, and why every test should be judged on revenue.
Basics
A/B tests compare whole versions of a page; multivariate tests change several elements at once in every combination. How they differ and when each is worth it.
The A/B testing mistakes that produce fake winners: stopping early, tiny samples, the wrong metric and broken splits. What each costs and how to fix it.
The ten steps of a trustworthy A/B test, from picking the metric and sizing the sample to checking the split, reading the result and keeping a holdout.
Split testing pays when the revenue through tested pages is big enough that a few real lifts beat tool and time costs. Here's the formula and a worked example.
AI is good at test ideas, writing variations and spotting funnel leaks. It shouldn't call winners. Use plain statistics on real visitor data for that decision.
A/B testing shows two versions of a page to random halves of your visitors and measures which one earns more. How it works, and when it is worth doing.
Statistics
Frequentist tests give a p-value; Bayesian tests give the chance B beats A. How the two differ, the same data analyzed both ways, and why neither fixes peeking.
A holdout keeps a small random share of visitors on the old version after you ship winners, so you can measure what those winners really earned.
Run an A/B test until it reaches its planned sample size, rounded up to whole weeks. That means at least one week and usually no more than four to six.
The A/B test sample size formula, a worked example and a lookup table. At a 3% conversion rate, detecting a 20% lift needs about 13,900 visitors per version.
MDE is the smallest lift your A/B test can reliably detect with its traffic. Learn the formula, a worked example, and how to pick an MDE before you launch.
What statistical significance and p-values mean in an A/B test, what 95% does and doesn't tell you, and how often significant winners turn out to be false.
Sample ratio mismatch means your A/B test's traffic split doesn't match what you set. How to check with a chi-square test, worked examples, causes and fixes.
Checking an A/B test daily and stopping at the first significant result can turn a 5% false win rate into 25% or more. A simulation, sources and fixes.
What to test
What to A/B test in checkout for e-commerce and SaaS, based on why shoppers abandon, how to judge tests on revenue per visitor, and how to run them safely.
What to A/B test on a call to action button: the copy and what happens after the click first, placement second, color last, with the traffic each test needs.
Write headlines that make different promises, split traffic evenly and judge on revenue, not clicks. What to try, how long to run it and what can go wrong.
Pick the activation event that predicts paying, track time to first value, and decide onboarding A/B tests on revenue measured 30 to 90 days later.
Start where your funnel loses the most money, then score ideas on impact, confidence and ease and check you have the traffic. A worked example with real maths.
Landing page A/B testing in the order that moves revenue: offer, headline, proof, call to action, then design, and the traffic each test needs.
Pricing and revenue
How to test showing annual or monthly billing first on your pricing page, why cash on day one misleads, and how to judge the result on 12-month revenue.
When a free trial earns more than charging up front, how to test trial vs no trial and trial length, and what app data says about short and long trials.
How to A/B test a subscription cancel flow with pause offers, downgrades and exit surveys, and why to judge it on revenue kept after 30 and 90 days.
How to A/B test pricing for SaaS and e-commerce: what to test, what is fair and legal to vary, the guardrails to set, and how to judge the result on revenue.
Tag each payment with the visitor's test version via client_reference_id or metadata, subtract refunds, then divide net Stripe revenue by visitors per version.
What to A/B test on a mobile app paywall, why to judge it on revenue per install instead of trial starts, weekly vs annual plans and the app store rules.
A version can win on signups and still lose money. Revenue per visitor counts conversion, price, plan mix and refunds, so it picks the version that earns more.
Channels
Send each subject line to a random slice of your list and pick the winner on clicks or revenue per recipient, not opens, which Apple Mail inflates since 2021.
Use Meta's A/B test tool, change one thing, run 7 to 30 days, and pick the winner on paying customers per dollar from your payment data, not on CTR.
Shopify's Rollouts can test themes, checkout, prices and discounts on the Grow plan and up. Here's what it allows, where it stops, and how to judge on revenue.
SEO and AI SEO
SEO split testing changes a random half of similar pages and compares their Google clicks with the unchanged half, because Googlebot can't see two versions.
Run a fixed panel of buyer prompts several times a week, track how often you're mentioned and cited, and test page changes against similar unchanged pages.
Low traffic
Let Outtest run your split tests
AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.