outtest
Get started
Statistics

How long should you run an A/B test?

Run an A/B test until it reaches its planned sample size, rounded up to whole weeks. That means at least one week and usually no more than four to six.

Updated 28 September 2026 · 6 min read

The short answer

Run an A/B test until it reaches the sample size you calculated before launch, then round up to whole weeks so every weekday is counted equally. Run at least one full week (two if returning visitors might react to novelty) and usually no more than four to six, because cookies expire, seasons shift and traffic mix drifts. Days needed = visitors needed across all versions / visitors entering the test per day.

Free tool: A/B test duration calculator. No signup.

Run an A/B test until it reaches the sample size you calculated before launch, then round up to whole weeks. In practice that means at least one full week and, for most sites, no more than four to six. The duration isn't a matter of taste: it falls out of your traffic, your conversion rate and the smallest lift you want to detect.

The formula is simple: days = (visitors needed per version × number of versions) / visitors entering the test per day. The /tools/ab-test-duration-calculator does it in one step.

A worked example

An online store's product pages get 1,500 visitors a day and convert 2.5% of them. The team wants to detect a 20% relative lift (2.5% to 3.0%) at 95% confidence and 80% power.

  1. Sample size: the standard formula gives 16,792 visitors per version (/guides/ab-test-sample-size shows the maths).
  2. Total for two versions: 2 × 16,792 = 33,584.
  3. Days: 33,584 / 1,500 = 22.4 days.
  4. Round up to whole weeks: 28 days, or four weeks.

The test runs for four weeks, and the team reads the result once, at the end.

Rounding 22.4 days up to 28 feels wasteful, but stopping at 22 or 23 days would count some weekdays three times and others four. If your Saturday shoppers convert differently from your Tuesday ones, the result would be tilted toward whichever days got the extra count.

Why whole weeks

Visitor behavior changes through the week. B2B sites are busy Monday to Friday and quiet at weekends. Many stores see a different audience on Sunday evening from Tuesday lunchtime. A test that doesn't cover each day equally measures a skewed slice of your customers.

In Controlled experiments on the web, Ronny Kohavi and colleagues write that even when you have enough visitors to finish in hours, they strongly recommend running experiments for at least a week or two, then continuing by multiples of a week so day-of-week effects can be analyzed.

Durations at a glance

Days needed (rounded up to whole weeks) at a 3% baseline conversion rate, 95% confidence, 80% power and a 50/50 split:

Visitors a day into the test 10% lift 15% lift 20% lift 30% lift
500 217 98 56 28
1,000 112 49 28 14
2,000 56 28 14 7
5,000 28 14 7 7
10,000 14 7 7 7

Anything above 42 days in this table is a test you probably shouldn't run as designed. At 1,000 visitors a day, a 10% lift takes 112 days to detect. Test a change that could plausibly move the number by 20% or more instead, or pick a page with more traffic or a higher conversion rate.

The minimum is one full week, sometimes two

Even with huge traffic, don't stop before seven days. Beyond the day-of-week effect, the first days of a test attract your most frequent visitors disproportionately, because they're the ones who show up first. A result based on them may not hold for everyone else.

Two weeks is safer when returning visitors might react to the change just because it's new. In Seven Rules of Thumb for Web Site Experimenters, Kohavi and colleagues recommend running experiments for two weeks and looking for novelty effects, where a lift fades as people get used to the change. They also note these effects are uncommon in practice, so the extra week is insurance, not a rule.

Outtest's Referee won't call a winner before a test has run 7 full days. The minimum can be set anywhere from 1 to 30 days in Settings.

The maximum is about four to six weeks

Long tests get less reliable, not more. Four things drift:

  • Cookies expire. Most testing tools remember which version a visitor saw in a cookie. Since Intelligent Tracking Prevention 2.1 in 2019, Safari caps persistent cookies created by JavaScript at seven days. A Safari visitor who returns after that can be assigned again and may see the other version, which blurs the difference between groups.
  • The season changes. A six-week test that spans Black Friday or a summer lull compares a mix of very different weeks.
  • Your traffic changes. New campaigns, a press mention or a search ranking shift change who arrives, and the early and late parts of the test measure different audiences.
  • You can't run the next test. Every week spent on a test that can't finish is a week not spent on one that can.

Outtest ends any test that hasn't cleared all three of its bars by day 42 as a draw, and the original stays. That's a reasonable default for most sites.

Business cycles

Whole weeks are the minimum cycle. Some businesses have longer ones that the test should cover:

  • Purchase cycles. If buyers usually visit several times over two weeks before paying, a one-week test mostly measures people who were already close to buying before it started. Run at least one full purchase cycle.
  • Pay cycles. Consumer spending often changes around paydays. A test of a pricing page might need to span a month-end.
  • Billing cycles. For subscription tests where the result is a renewal or a first payment after a trial, the test must run long enough for that event to happen. A 14-day trial test needs at least 14 days after the last visitor enters, plus time for payments to settle. /guides/free-trial-vs-no-trial covers this.
  • Campaigns and holidays. Avoid starting a test the day a big sale begins, or make sure both versions get the same sale for the full run.

Don't stop early, and don't extend

Two tempting moves both break the maths:

  • Stopping when the result looks significant. A test with no real difference crosses the 95% line at some point far more often than 5% of the time if you check daily. Evan Miller showed that checking after every visitor pushes the real false positive rate to 26.1% (How Not To Run an A/B Test).
  • Extending a finished test because it's "nearly there". Adding days until the result crosses the line is the same problem in reverse. You're giving chance extra tries.

/guides/peeking-problem-ab-testing shows a simulation of both. If you need the option to stop early, use a method designed for it, such as a sequential test.

It's fine to stop early if a version is broken or clearly losing money. Protecting revenue matters more than a clean experiment.

When the answer is "too long"

If the calculator says your test needs 90 days, you have four options:

  1. Test a bigger change. The sample size falls with the square of the lift, so sizing for a 20% lift instead of 10% cuts the duration by about three quarters.
  2. Test where conversion rates are higher. A checkout step converting 50% needs a small fraction of the visitors a homepage converting 2% does (/guides/what-to-ab-test-first has an example).
  3. Accept a larger minimum detectable effect. The /tools/minimum-detectable-effect-calculator tells you what lift your traffic can detect in, say, four weeks. /guides/minimum-detectable-effect explains how to use it.
  4. Don't A/B test that decision. Make the change based on customer research and watch the numbers. /guides/ab-testing-low-traffic covers this.

A quick checklist

  • Sample size calculated before launch from the page's real traffic and conversion rate.
  • Duration rounded up to whole weeks, at least one.
  • Test spans at least one full purchase or billing cycle.
  • Test ends within about six weeks, or it's redesigned.
  • End date written down, and nobody calls a winner before it.

Questions people ask

Can I stop an A/B test early if it's winning?+

Not safely, unless you planned for it with a sequential method. Significance swings in the early days, and stopping at the first good reading makes false winners several times more likely than the 5% you think you have. You can stop early if a version is broken or losing badly; just don't declare a winner before the planned end.

What is the minimum time to run an A/B test?+

One full week, even if you reach the sample size sooner. Weekday and weekend visitors often behave differently, and a test that misses some days measures a different audience from the one you'll ship to. Two weeks is safer when returning visitors might react to the change simply because it's new.

What is the maximum time to run an A/B test?+

Usually four to six weeks. Longer tests suffer from cookie expiry (Safari caps cookies set by JavaScript at seven days), seasonal shifts and changing traffic, and they block other tests. If the maths says you need three months, test a bigger change or a busier page instead.

Should I keep running a test that's almost significant?+

Only if it hasn't reached its planned sample size. Extending a finished test until it crosses the line is another form of peeking and raises the false win rate. If it ends short of significance, call it a draw and move on.

Read next

Let Outtest run your split tests

AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.