Why you should judge A/B tests on revenue per visitor
A version can win on signups and still lose money. Revenue per visitor counts conversion, price, plan mix and refunds, so it picks the version that earns more.
Updated 28 September 2026 · 7 min read
Judge A/B tests on revenue per visitor: the money a version brought in, after refunds, divided by every visitor who saw it. It catches the common case where a version wins more signups by selling cheaper plans and earns less overall. Revenue is noisier than a conversion rate, so a revenue test can need several times more visitors unless you cap unusually large payments at a limit set before the test starts.
Judge an A/B test on revenue per visitor, meaning the revenue each version brought in, after refunds, divided by all the visitors who saw it. Conversion rate counts how many people did something. Revenue per visitor counts the money that came with them, so it catches the version that wins more signups by pushing people to a cheaper plan, a bigger discount or a lower price.
The cost is noise. Most visitors pay nothing and a few pay a lot, so revenue per visitor needs more traffic than a conversion rate to reach the same confidence. You handle that by sizing the test on revenue, capping unusually large payments at a limit you pick before the test starts, and counting revenue over a window long enough to include refunds and the first renewal.
What revenue per visitor measures
Revenue per visitor = (revenue from visitors in a version, minus refunds) divided by the visitors in that version.
Three details matter:
- Every visitor counts, including the ones who paid nothing. That is the difference from average order value, which only looks at buyers.
- It equals conversion rate times revenue per buyer. A 10% drop in conversion with a 20% rise in revenue per buyer is a net gain of 8%, because 0.9 × 1.2 = 1.08.
- Revenue belongs to the version the visitor was first assigned to, even if they pay a week later on another device. That needs a stable visitor ID, covered below.
A worked example where more signups means less money
This is a made-up example with round numbers. A SaaS pricing page sells Starter at $29 a month and Pro at $79. Version A marks Pro as "most popular". Version B marks Starter and moves Pro to the right. Each version gets 10,000 visitors.
| Version A (Pro highlighted) | Version B (Starter highlighted) | |
|---|---|---|
| Visitors | 10,000 | 10,000 |
| Paying customers | 300 (3.0%) | 360 (3.6%) |
| Chose Pro at $79 | 120 (40%) | 54 (15%) |
| Chose Starter at $29 | 180 | 306 |
| First-month revenue | $14,700 | $13,140 |
| Revenue per visitor | $1.47 | $1.31 |
Here is the maths. Version A earns 120 × $79 + 180 × $29 = $9,480 + $5,220 = $14,700. Version B earns 54 × $79 + 306 × $29 = $4,266 + $8,874 = $13,140.
Version B wins 20% more customers and earns 10.6% less. A tool that judges on signups would call B the winner, and with some confidence. At 3.6% against 3.0% on 10,000 visitors each, B has about a 99% chance of having the higher conversion rate. Ship it and you lose $0.16 per visitor, and the gap grows every month those lower-paying customers stay subscribed.
Plug your own numbers into the revenue per visitor calculator to see where your versions land.
Why clicks and signups point the wrong way
This is not a rare edge case. Two examples from a paper by Ron Kohavi and colleagues on rules of thumb for online experiments, drawn from thousands of online experiments:
- At Amazon, a credit card offer kept winning the top slot on the home page even though very few people clicked it. Each signup was so profitable that its expected value beat everything else. Moving it to the shopping cart page was worth tens of millions of dollars in profit a year.
- Microsoft Office Online tested a page redesign that showed the product's price. Clicks per user fell 64%. The team had assumed clicks times a fixed conversion rate equals revenue, but the people who still clicked were better qualified and bought at a much higher rate.
The patterns behind these show up on small sites too:
- A discount raises conversion and lowers revenue per buyer.
- A free plan raises signups and can cut paid conversion.
- A bundle or upsell lowers conversion slightly and raises order value.
- An annual plan pulls a year of revenue into the first payment.
In each case, one number goes up and another goes down, and only revenue per visitor shows the net result.
How variance works for revenue metrics
A conversion is yes or no. Its variance per visitor is p × (1 − p), where p is the conversion rate. At 3%, that is 0.03 × 0.97 = 0.0291.
Revenue per visitor is mostly zeros with a few large numbers. Its variance is the average of each visitor's revenue squared, minus the square of the average. Large payments count heavily because they get squared.
A standard shortcut for sample size, at 80% power and 5% significance, is:
Visitors per version ≈ 16 × variance ÷ (difference you want to detect)²
Here is what that gives for a 10% lift, using the pricing page example above.
| Metric | Average per visitor | Variance | Visitors per version for a 10% lift |
|---|---|---|---|
| Conversion rate | 3.0% | 0.0291 | about 52,000 |
| Revenue, two monthly plans | $1.47 | 87.9 | about 65,000 |
| Revenue, plus a $790 annual plan bought by 0.2% of visitors | $2.89 | 1,317 | about 252,000 |
| Same, with the annual payment capped at $300 | $1.91 | 254 | about 111,000 |
For the second row, 1.8% of visitors pay $29 and 1.2% pay $79. The average squared revenue is 0.018 × 29² + 0.012 × 79² = 15.14 + 74.89 = 90.03. Subtract 1.47² = 2.16 to get a variance of 87.87. Then 16 × 87.87 ÷ 0.147² ≈ 65,000.
Two monthly prices barely hurt, since they only need about 25% more visitors than the conversion rate. Adding one rare, large payment makes the test nearly five times bigger. That is the usual shape of revenue data. A few big orders, annual plans or bulk buyers do most of the damage.
Real data looks worse than this example. In the Kohavi paper, Bing's revenue per user was so skewed that their rule of thumb needed about 114,000 users per version before the average behaved normally. Capping revenue at $10 per user per week cut that to about 9,700, and with the same sample the capped metric could detect a change 30% smaller.
Capping large payments
Capping (statisticians call it winsorizing) means treating any visitor's revenue above a limit as the limit. It keeps every visitor in the test and stops one huge order from deciding the result.
- Pick the cap before the test starts. A common choice is the 99th percentile of revenue per paying visitor over the last 90 days.
- Apply the same cap to both versions.
- Report the uncapped result as well. If the two disagree, look at the largest orders by hand before you decide.
- Never cap by deleting orders. Deleting orders changes the answer itself.
Which revenue to count
Decide these before launch and write them down:
- Net revenue. Subtract refunds and chargebacks, and leave out sales tax and VAT.
- A fixed window per visitor. Count payments made within, say, 30 days of each visitor's first visit. Without a fixed window, visitors from the first day of the test get more time to pay than visitors from the last day.
- A longer look for subscriptions. First-month revenue misses churn. Check again at 90 days, because a cheaper or discounted version often keeps customers for a different length of time.
- Trials. If there is a free trial, the window has to cover the trial plus the first payment, or you are measuring trial starts with extra steps.
- One currency. Convert at a fixed rate for the whole test so exchange rates don't pick the winner.
After a winner ships, a holdout group that keeps seeing the old version shows whether the lift held up over months.
How to track revenue per visitor
- Give each visitor a stable ID and a version, stored in a first-party cookie, so they see the same version on every visit.
- Pass the ID and version to your payment tool when someone pays. In Stripe, that means adding them as metadata on the checkout session or customer. The Stripe testing guide covers the details.
- Pull payments and refunds for the window, join them to visitors on the ID, sum per version and divide by visitors.
- Check that the traffic split is what you set with the sample ratio mismatch checker. A broken split makes any metric wrong.
Outtest does this join for you. It reads payments read-only from Stripe, Shopify, Paddle and other payment tools, judges every website, pricing and ad test on revenue per visitor, and shows signups only as context.
Tests where revenue per visitor and conversion rate often disagree
These are the tests most likely to fool a signup metric, and what to measure on each:
- The plan marked "most popular". Measure revenue per visitor and the share of buyers on each plan.
- Showing or hiding a free plan. Measure revenue per visitor at 30 days and the share of free users who upgrade.
- A first-order discount banner. Measure revenue per visitor, average order value and margin after the discount.
- A higher free shipping threshold. Measure revenue per visitor and average order value.
- Annual billing selected by default. Measure revenue per visitor at 30 and 90 days and the share choosing annual.
- An order bump or upsell at checkout. Measure revenue per visitor and checkout completion.
- Per-seat price shown instead of a total. Measure revenue per visitor and seats per purchase.
For each, set the sample size from revenue variance, not from conversion rate. The sample size guide explains the inputs, and statistical significance explained covers how to read the result.
When a signup metric is still fine
If you have no payment data, such as lead generation where sales close offline weeks later, you have to judge on qualified leads. Do it knowingly, and check the winner against closed revenue once it comes in. Anywhere you can see payments, use them. A test is only worth running if it changes how much money you make, so measure that.
Questions people ask
What is revenue per visitor in A/B testing?+
Revenue per visitor is the total revenue a version produced, after refunds, divided by the number of visitors assigned to it. Visitors who bought nothing count as zero. It combines conversion rate and revenue per buyer into one number, so it shows which version made more money per person who saw it.
Is revenue per visitor better than conversion rate?+
For picking a winner, yes, whenever the change can affect what people pay: pricing pages, plan layouts, offers, upsells and checkout. Conversion rate is still useful as context and as an early signal. When the two disagree, revenue per visitor is the one that matches your bank balance.
Why do revenue tests need more traffic?+
Most visitors pay nothing and a few pay a lot, so revenue varies far more from person to person than a yes or no conversion does. More variation means you need more visitors to tell a real difference from noise. Capping very large payments at a limit you set in advance removes much of that variation.
How do I calculate revenue per visitor for a subscription?+
Pick a fixed window, such as 30 or 90 days after each visitor's first visit, and add up the payments each visitor made in that window, minus refunds. Divide by the visitors in that version. Use the same window for both versions so neither gets extra time to pay.
Read next
Let Outtest run your split tests
AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.