Experimentation / CRO

Experimentation and Testing Programs acknowledge that the future is uncertain. These programs focus on getting better data to product and marketing teams to make better decisions.

Research & Strategy

We believe that research is an integral part of experimentation. Our research projects aim to identify optimization opportunities by uncovering what really matters to your website users and customers.

Data and Analytics

90% of the analytics setups we’ve seen are critically flawed. Our data analytics audit services give you the confidence to make better decisions with data you can trust.

Can You Run Multiple A/B Tests at the Same Time?

Cover image for the Speero blog post Can You Run Multiple A/B Tests at the Same Time, with the title beside the Speero logo and a graphic of two overlapping experiment traffic streams.
The Diagnosis: Yes, and the fear costs more than the problem. Microsoft looked at every pair of A/B tests running on the same day across four products, each running hundreds of tests daily on millions of users. Three of the four found no interaction effects at all. The fourth found abnormal signals in 0.002% of test-pair metrics, or 1 in 50,000.
Meanwhile the fix has a price you can calculate. Splitting traffic between two mutually exclusive tests raises each one's minimum detectable effect by 41%, which is enough to move a page from "can referee a test" to "cannot."
So the real question isn't whether interactions happen. It's whether your surface can afford the insurance.

You're running a test on the product page. Someone else wants to run one in checkout. The question lands in Slack and stalls there for a week.

Here's the answer, and then the arithmetic behind it.

Run them concurrently. You can run multiple A/B tests at the same time because interaction effects are rare enough that the cost of preventing them usually exceeds the cost of ignoring them. The exception is narrow and worth knowing, and it depends on your traffic rather than on your caution.

"Overlapping experiments are the least of several evils." Lukas Vermeer, Director of Experimentation at Vista and previously Director of Experimentation at Booking.com.

Four years after he said it to us, the published evidence has caught up.

What an interaction effect actually is

An interaction effect happens when the combined impact of two changes differs from the sum of their individual impacts.

Your product page test lifts add-to-cart by 3%. Your checkout test lifts completion by 2%. If both are live and the true combined effect is 5%, there's no interaction. If it's 1%, or 9%, the two treatments are talking to each other.

That's the whole concern. Notice what it isn't. Two tests running at once do not "pollute" each other's data by default.

Randomization puts roughly half of each test's traffic into each of the other test's arms, so the other experiment becomes noise distributed evenly across your variants rather than bias pushing in one direction.

The worry is specifically about the cases where that even distribution breaks down.

How rare is rare

For four years this post asserted that interaction effects are rare without a number behind it. Here's the number.

Monwhea Jeng at Microsoft's Experimentation Platform ran a meta-analysis of A/B interactions across four major products, each running hundreds of A/B tests per day on millions of users. The method was to take every pair of tests running on the same day and check the distribution of interaction p-values.

Three of the four products found none. Every result was consistent with no interaction at all. Across all four, the p-value distribution came back very close to uniform, which is what you'd expect if interactions simply aren't there.

The fourth product found abnormally small p-values in 0.002% of test-pair metrics, or 1 in 50,000.

One detail matters more than the headline figure. In the cases where something did show up, there were no instances of two treatment effects being both statistically significant and moving in opposite directions. The scenario people actually fear, where two winners combine into a loser, didn't appear.

"The vast majority of A/B tests either don't interact or have only relatively weak interactions." Monwhea Jeng, Principal Data Scientist, Microsoft
How often concurrent A/B tests actually interact Chart showing Microsoft's analysis of concurrent A/B test pairs across four products: three products found no evidence of interaction effects at all, and the fourth found abnormal signals in 0.002% of test-pair metrics, or 1 in 50,000. Every pair of concurrent A/B tests, four products, one day Each product runs hundreds of A/B tests per day on millions of users Share of test-pair metrics showing an abnormal interaction signal Product A None found Product B None found Product C None found Product D 0.002% of pairs 1 in 50,000 the highest rate any product recorded Across all four products the p-value distribution came back close to uniform. No case showed two significant effects moving in opposite directions. Source: Monwhea Jeng, Microsoft Experimentation Platform, A/B Interactions: A Call to Relax. speero.com

Georgi Georgiev reached a compatible verdict from a different direction. His 2017 simulation of concurrent tests used only large-uplift winners, which is the most favorable possible condition for producing an interaction. He still concluded that "there is no certain way to establish the likelihood of harmful interference between concurrent A/B tests, nor the impact of such events."

Read those two together and the picture is consistent. Interactions are hard to predict in advance and rare enough in practice that predicting them isn't where your effort belongs.

What isolating traffic costs you

The reason this matters is that the cure has a price, and almost nobody quotes it.

Minimum detectable effect scales with the inverse square root of your sample size. Halve the traffic going into a test and its MDE rises by the square root of 2, which is 41%.

That's the number to carry into the next planning meeting. A page that could detect a 10% relative lift detects 14.1% once you make its test mutually exclusive with another. A page at 20% moves to 28%.

Microsoft's team quantified the same relationship from the other end. In their KDD paper on running controlled experiments at large scale, Kohavi and colleagues put a figure on it. Improving sensitivity by a factor of 10, from a 5% detectable delta to 0.5%, requires 100 times more users. Sensitivity is expensive, and isolation spends it on insurance against a 1-in-50,000 event.

Work a concrete case. A demo-request page converting at 3%, and you want to detect a 10% relative lift at 95% confidence and 80% power. Using the standard rule-of-16 approximation you need roughly 51,700 visitors per variant, so about 103,000 to call the test. You can check that with the A/B test calculator.

At 16,000 visits a month, one test using all the traffic resolves in about 6.5 months.

Now run two tests as mutually exclusive groups. Each gets 8,000 visits a month, so each takes about 13 months. Sequential isolation gives you the same total: 6.5 months for the first, 6.5 for the second.

Both isolation strategies take twice as long to produce two answers as simply letting the tests overlap. That's the trade you're making, and it's the same trade whether you queue the tests or split the traffic.

What isolating traffic costs your detectable effect Chart showing the cost of isolating traffic between two A/B tests. Splitting traffic in half raises the minimum detectable effect by 41 percent. A Tier 1 surface detecting a 10 percent lift moves to 14.1 percent, which is Tier 2. A Tier 2 surface detecting 20 percent moves to 28.3 percent, which is Tier 3 and can no longer referee a test. Isolating traffic costs 41% of your detectable effect Smallest relative lift a surface can detect, before and after splitting its traffic with a second test TIER 1 TIER 2 TIER 3 · CANNOT REFEREE A TEST Tier 1 surface still readable, one tier worse 10% 14.1% Tier 2 surface drops out of testable range 20% 28.3% 0% 5% 10% 15% 20% 25% 30% Full traffic Traffic split with a second test Detectable effect scales with 1 over the square root of sample size, so half the traffic costs you the square root of 2 speero.com

The answer changes by tier

A universal answer to "can I overlap" is the wrong shape. Whether you can afford isolation depends on how much detectable effect you have to spend, and that varies enormously by surface.

Speero tiers surfaces by the smallest relative lift they can detect in a fixed window. For lower-traffic B2B and SaaS on a 4-week window, Tier 1 detects 10% or better, Tier 2 sits between 10% and 20%, and Tier 3 is everything above. The test bandwidth calculator will tell you which tier a given page is in.

Now apply the 41% penalty. A Tier 1 surface becomes Tier 2. A Tier 2 surface becomes Tier 3, which by definition can no longer referee a normal test.

On a high-traffic ecommerce page, isolation is a scheduling inconvenience. On a low-traffic B2B page, isolation converts a readable question into an unreadable one.

That's not hypothetical. In B2B programs we run, plenty of tests resolve on single-digit form submits per variant. Reading a result from 8 submissions against 4 is already at the edge of what the numbers support. Halving the traffic into that test doesn't weaken the read. It removes it.

So the rule that survives contact with real programs: the lower your traffic, the less you can afford to isolate, which is the opposite of how most teams behave. Low-traffic programs tend to be the most cautious about overlap, and they're the ones for whom caution is most expensive.

Two experiments, one funnel

Three Ways to Run Them, and What Each One Costs

A page at 16,000 visits a month, 3% baseline, detecting a 10% relative lift.

Yellow marks the option that costs you least on each row.
  OverlappingBoth on full traffic Parallel isolationMutually exclusive groups Sequential isolationOne test, then the other
Time to both results 6.5 months 13 months 13 months
Detectable effect per test 10% 14.1% 10%
Learn the interaction Yes, if you check No No, needs a third test
Tier cost None Drops one tier None
Right when Almost always Same page, same goal, expensive decision One test is blocking a decision

Swipe to see all columns →

Time to result assumes the rule-of-16 approximation: roughly 51,700 visitors per variant to detect a 10% relative lift at a 3% baseline, 95% confidence and 80% power. Parallel isolation halves the traffic reaching each test, which raises its minimum detectable effect by the square root of 2.

How to detect an interaction if you want to check

Overlapping doesn't mean flying blind. If you want to know whether two tests interacted, you can look, and looking after the fact is cheaper than preventing it beforehand.

Segment one test's results by the other test's variant assignment. It's the same operation as splitting by device or traffic source: you're just using experiment membership as the segment. If the effect of test A is materially different among test B's control than among test B's treatment, you have a candidate interaction.

Then apply a correction, because you're now running many comparisons. Georgiev's recommendation stands: use a Šidák correction for a handful of interaction checks, and if you're testing interactions in the hundreds or more, use the Benjamini-Hochberg-Yekutieli false discovery rate adjustment instead.

For a quick read before you commit to any of that, Lukas Vermeer's XY calculator estimates the chance of overlap mattering given your traffic and conversion numbers.

You don't have to master the corrections to make a decision.

"Generally, these are rigorous ways to address the interaction effects. But also just knowing these effects do hurt detectability level, and if you and your team are okay with this tradeoff, then you take the hit on confidence and move forward."

That's Ben Labay, CEO of Speero.

One caveat on all of this. Interaction checks are underpowered almost by construction, because detecting an interaction reliably needs far more traffic than detecting either main effect. A clean interaction check is a luxury of high-traffic programs. Everyone else is making a judgment call, and should say so rather than pretending the segment split settled it.

What the interaction check can actually see

That last caveat deserves a number, because the number changes what the check is for.

An interaction contrast is the difference of two differences. You are comparing how test A performed inside test B's control against how it performed inside test B's treatment, which means four cells of traffic contribute variance instead of two. Run the arithmetic and it lands somewhere uncomfortable. Detecting an interaction the same size as a main effect takes 4 times the total traffic.

The sharper version of that is worth carrying into the conversation.

Take the page from earlier, 3% baseline, powered to detect a 10% relative lift, about 101,500 visitors in total. Overlap two tests on it and that traffic splits across four combinations of assignments, roughly 25,400 each. At that cell size the smallest interaction you can detect is 20% relative. Double the main effect you were originally looking for.

So the segment split is not a clean bill of health. It rules out an interaction twice the size of your treatment effect, and it is silent on anything smaller.

Which is fine, because that is the case worth ruling out. The scenario people describe when they resist overlapping tests is the catastrophic one, two winners combining into a loser, and a catastrophic interaction is large by definition. Jeng's meta-analysis found no instances of two significant effects moving in opposite directions across four products running hundreds of tests a day. The check screens for the failure mode that would actually hurt you, and the traffic you have is enough for that.

The honest thing is to say which question you answered. "We segmented by the other test's assignment and saw nothing" invites the reader to hear "no interaction." What you can defend is narrower. You saw no interaction large enough to be visible at your sample size, and you should write down what that size was.

When the decision really does ride on a small effect and the check cannot reach it, stop analyzing and start replicating. Three methods for confirming test effects covers the options. A back test on the winner you are about to build a roadmap around costs less than the traffic you would need to power the interaction properly. Trustworthy Online Controlled Experiments is the reference if someone wants to argue the math.

The false positive nobody is worried about

Here is the part that should worry you when you run multiple A/B tests at the same time, and it has nothing to do with the tests talking to each other.

Every test you run at 95% confidence carries a 5% chance of a false positive on its own. That is the deal you signed. Run a lot of them at once and the arithmetic compounds in a way that is easy to state and rarely stated.

At 5 concurrent tests, the chance that at least one reports a win that isn't there is 22.6%. At 10 it's 40.1%. At 20 it's 64.2%. At 50 it's 92.3%. A program running 20 tests a month expects 1 false positive a month with no real effects anywhere in the portfolio.

Now put the two risks on the same scale, generously.

Twenty concurrent tests give you 190 test pairs. Check 10 metrics on every pair, which is more diligence than anyone actually applies, and you have 1,900 test-pair metrics. At Jeng's rate of 1 in 50,000 you expect 0.04 abnormal interaction signals. In the same month, alpha alone hands you 1.0 false positives.

The concurrency people fear is roughly 26 times more likely to give you a fake winner than a real interaction. Same month, same program, same traffic.

The response to this is not a correction applied across your whole program. Family-wise corrections belong inside a test, across the variants competing for one decision. That is why Bonferroni shows up when you run A against B and C. Twenty unrelated tests on twenty unrelated pages are twenty separate decisions, and correcting across them would just make every individual answer harder to reach.

What the arithmetic argues for is a confirmation habit on the results you intend to build on. Ship the small wins and move on. Back test the one that is about to justify a quarter of roadmap.

It also argues for watching your win rate at program level. Microsoft's number is that roughly a third of tested ideas improve the metric they targeted. A program reporting a 60% win rate across a busy concurrent portfolio is not outperforming Microsoft. It is looking at alpha. Program metrics is where that gets tracked rather than assumed.

And there is one failure mode that beats both of these for frequency. A sample ratio mismatch means the assignment or the logging broke, and a broken split produces confident nonsense that no interaction analysis will catch. Check it first, on every test, before you look at anything else. The KDD taxonomy of SRM causes, built from four software companies and more than 25 products, is the reference for tracing one.

Decide the action before you launch

Here's the step this post skipped for four years, and the one that actually determines whether any of the above is useful.

Write down what you'll do with each possible outcome before either test goes live.

If test A wins and no interaction shows up, you ship A. Fine. But what if A wins, B wins, and the segment split suggests they cancel? What if A wins only among B's control group? Decide those in advance and the answer is a decision. Decide them afterwards and the answer is whatever the person with the strongest opinion wanted anyway.

Speero's experimentation decision matrix exists for exactly this: define the action for every combination of primary and guardrail outcomes, then lock it at launch. The test phase gate framework is where that commitment gets governed rather than remembered.

A worked pre-commitment for two overlapping tests, as an example rather than a template:

Ship both if neither shows an interaction signal. Ship the one with the larger effect and re-run the other alone if the segment split disagrees by more than the correction threshold. Ship neither and run a combined variant if both look positive alone but negative together.

Three rules, agreed before launch, and the post-test meeting takes ten minutes instead of two weeks.

When Not to Run Multiple A/B Tests at the Same Time

There is a real case for it. Two tests on the same page, targeting the same goal, where a false read would be expensive. Pricing is the usual example.

You have three options, in rough order of how much they cost you.

Add a variant instead of adding a test. If both ideas live on the same page and serve the same hypothesis, make the second idea a third arm. Instead of A and B, run A, B, and C, where C is B plus the change you wanted to test separately. You pay in traffic for the extra arm and you get the interaction for free, because it's now a factorial design rather than two experiments hoping not to collide.

Use mutually exclusive groups. Optimizely supports mutual exclusion across Web Experimentation, Personalization, Performance Edge, Feature Experimentation and Full Stack. One operational constraint is worth knowing. Experiments have to be in the group from the start of the first experiment to the end of the last, to keep bucketing fixed. Their own guidance is the same as this post's. Make experiments mutually exclusive only when necessary, to avoid unduly restricting traffic. VWO offers the same capability, now documented under Wingify, as an enterprise feature on FullStack with a limit of 10 campaigns per group.

Sequence them. Run one, then the other. It costs the same total time as parallel isolation, and it's the right answer when one test is blocking a decision and the other can wait. The should I run an A/B test checklist is a better place to start than the queue, since some of what's waiting shouldn't be tested at all.

One thing that will wreck a concurrent test far more reliably than any interaction: a broken traffic split. We've had tests come back unreadable because allocation drifted between variants, and no amount of interaction analysis rescues that. Check your sample ratio before you check for interactions. The A/B testing QA process catches this class of problem before launch, which is the only place it's cheap to catch.

Why the question stalls in Slack for a week

The statistics in this post were settled by the third paragraph. The question still stalls, and it stalls for a structural reason rather than a mathematical one.

Nobody owns the answer.

Speero's work with Nils Stotz and Paul Drews at Leuphana University, and with Lukas Vermeer of Vista, sets out four mechanical preconditions for a team owning its own experimentation. A taxonomy for structuring experimentation teams has all four. The third one is the one that decides whether you can run multiple A/B tests at the same time without asking permission. Can a team change its own spec without approval from outside the team? Where the answer is no, every overlap question routes to somebody who is not accountable for the delay it causes, and that is exactly the condition under which a week disappears.

Price the delay and the asymmetry gets obvious.

That page doing 16,000 visits a month sheds about 3,700 visitors a week. Against the 103,000 the test needs, a week of stalling costs 3.6% of the answer, paid with certainty, on both tests, to insure against something running at 1 in 50,000. Do that four times a quarter and you have spent a seventh of a test on a conversation.

The reason it keeps happening is that the delay is invisible and the interaction is vivid. Nobody files a ticket for traffic that didn't get used.

Two structural fixes work, and they pull in opposite directions.

Give the decision an owner. A RASCI matrix for experimentation is where most programs find out they never wrote down who decides this, and naming one person is usually enough to turn a week into an hour.

Or remove the decision. Write the default into the program's operating rules so nobody has to ask. Overlap by default, isolate when two tests share a page and a goal and the decision is expensive. That sentence is the whole policy, and it belongs in a document rather than in a Slack thread that gets relitigated every quarter.

Centralizing the answer is the third option, and it scales worse than it looks. Meta managed shared test slots so teams could not launch freely and risk corrupting each other's data. That works at a volume where hundreds of tests run daily and a real collision is worth policing. At 20 tests a month the same mechanism just re-creates the queue this post is trying to dissolve. That is the pattern a center of excellence turning into a gatekeeper follows in general. The experimentation org structures blueprint covers which shape can hold a standing default and which one needs a referee.

The test of whether you have fixed this is short. Nobody should need permission to run multiple A/B tests at the same time on two unrelated pages. Ask what happens the next time two teams want the same page in the same week. If the answer is a conversation, you have a governance gap. If the answer is a rule, you have a program.

False positives against interaction signals as concurrent test count rises A line chart comparing two risks in a concurrent A/B testing program, both measured as expected occurrences per month. The first line, expected false positives, rises linearly at 0.05 per test, reaching 1.0 at 20 concurrent tests and 3.0 at 60. The second line, expected interaction signals, rises quadratically but stays almost flat along the bottom of the chart, reaching 0.04 at 20 concurrent tests and 0.35 at 60. At 20 concurrent tests a program expects 1.0 false positives and 0.04 interaction signals, a factor of 26. False positives assume independent tests at 95% confidence. Interaction signals assume every pair of concurrent tests is checked on 10 metrics at the observed rate of 1 abnormal signal per 50,000 test-pair metrics reported by Monwhea Jeng at Microsoft. The two lines only cross at about 501 concurrent tests. The practical conclusion is that in a normal program the false positive rate from alpha is a far larger source of wrong decisions than interaction between concurrent tests. TWO RISKS, SAME SCALE Where the Wrong Decisions Actually Come From Expected occurrences per month as a program runs more tests at the same time. 0 0.5 1.0 1.5 2.0 2.5 3.0 0 10 20 30 40 50 60 tests running at the same time false positives 3.0 a month at 60 interaction signals 0.35 a month at 60 At 20 tests at once: 1.0 against 0.04 26 times more fake winners than interactions False positives assume independent tests at 95% confidence, so 0.05 expected per test. Interaction signals assume every pair is checked on 10 metrics, at the rate of 1 abnormal signal per 50,000 test-pair metrics reported by Monwhea Jeng at Microsoft. The two lines only cross at about 501 concurrent tests. speero.com

Run Multiple A/B Tests at the Same Time by Default

The fear of interaction effects costs most programs more than interaction effects do.

The published evidence puts the rate at roughly 1 in 50,000 test-pair metrics, with three of four products at Microsoft finding none at all. The insurance against that costs 41% of your detectable effect, which on a low-traffic page is the difference between a test you can read and one you can't.

So the default is to run multiple A/B tests at the same time and check afterwards. Check for interactions afterwards when the stakes justify it, isolate only when two tests share a page and a goal and the decision is expensive, and write the decision rules down before anything launches.

Every test you don't run because you were waiting for a slot is a certain cost paid against an unlikely one. The should I run experiments simultaneously blueprint is the one-page version of this decision if you want something to put in front of a stakeholder.

Should I run experiments simultaneously? Speero blueprint Speero decision blueprint for running experiments simultaneously. The default answer is yes, always, with two options: complete time overlap or partial time overlap. Looking for exceptions, the first question is whether you test on a high traffic website with enough traffic to create mutually exclusive groups. If yes, and you have a tool with an exclusion or non-overlay feature, run the tests mutually exclusive and make sure you have enough statistical power. If either answer is no, ask whether your tests are on the same pages with the same hypothesis and the same conversion end goal. If yes, you could create an extra variant in your test instead. If no, ask whether your tests span multiple pages with the same conversion goal and same hypothesis. If yes, consider a multi-page experiment. If no, you can make an exception and wait for other tests to finish, which isolates the tests and decreases your test velocity. Should I runexperimentssimultaneously? Yes,always OPTION 1 Test with completetime overlap OPTION 2 Test with partialtime overlap LOOKING FOR EXCEPTIONS? Do you test on a high traffic website anddo you have enough traffic to createmutually exclusive groups? Yes Do you have access to a toolthat has an Exclusion ornon-overlay feature? Yes RUN TESTS MUTUALLYEXCLUSIVE. MAKE SUREYOU HAVE ENOUGHSTATISTICAL POWER. No No Do your tests check the following boxes: on the same page(s) same hypothesis same (conversion) end goal Yes YOU COULD CREATE AN EXTRAVARIANT IN YOUR TEST. No Do your tests check the following boxes: multiple pages same (conversion) goal same hypothesis Yes YOU COULD CONSIDER AMULTI-PAGE EXPERIMENT. No You can make an exception and wait for othertests to finish (isolate tests). Thisdecreases your test velocity.

If your program is queuing tests because the overlap question never got settled, that's a governance gap rather than a statistics problem. Speero's experimentation maturity audit is built to find that class of gap.

Did you like this article?

(Your feedback helps us write better!) worst 1 - 10 best

Did the article resonate with you?

What aspects did you enjoy or find lacking?
Were there elements you felt we should've covered?
Thank you.
Oops! Something went wrong while submitting the form.

Related Posts

Who's currently reading The Experimental Revolution?