Experimentation / CRO

Experimentation and Testing Programs acknowledge that the future is uncertain. These programs focus on getting better data to product and marketing teams to make better decisions.

Research & Strategy

We believe that research is an integral part of experimentation. Our research projects aim to identify optimization opportunities by uncovering what really matters to your website users and customers.

Data and Analytics

90% of the analytics setups we’ve seen are critically flawed. Our data analytics audit services give you the confidence to make better decisions with data you can trust.

Decision Velocity: Why One Confidence Bar Fails

Cover image for the Speero blog post Decision Velocity: Why One Confidence Bar Fails, showing a decision matrix diagram plotting reversibility against impact and risk, alongside a photo of author Paul Randall, Associate Director of Strategy at Speero.
The Diagnosis: Experimentation programs don't lose velocity because they move too fast. They lose it because every decision, a CTA colour and a pricing model alike, gets forced through the same 95% confidence bar. Matching that bar to the actual risk fixes it, but only if the thresholds are set before anyone sees the data. Skip that step and you're p-hacking on a shorter clock, not moving faster.

Most experimentation programs have a confidence problem. Not the one people usually blame.

The instinct is to assume the risk is moving too fast: shipping a bad variant, trusting a result before it's ready. So programs set one bar. 95% statistical confidence. Applied to everything. A below-the-fold copy tweak gets the same treatment as a new pricing model.

I've watched this play out enough times to know where it leads. The program feels rigorous. It also moves at a crawl. Tests queue up waiting for sample sizes they don't need. A pricing test sits behind a headline test in the same queue.

And every week a result sits at 88% confidence while the team holds out for 95, there's a real cost attached to that wait, one that almost nobody calculates and even fewer use to argue for moving faster.

That's the cost of delay. It's one of the most underused arguments in this industry.

Decision Velocity is what I call the discipline of matching your rigour to the actual risk of the decision. Done right, it puts your standards where they belong: on the decisions that can actually hurt you.

Here's how I build it.

The Real Cost of One-Size-Fits-All Confidence

A copy change below the fold is not a pricing model. The stakes are different. So is the reversibility. Most programs run them through the same process anyway.

2 problems follow from that.

First, traffic waste. Every session spent chasing an unnecessary confidence threshold on a low-impact tweak is a session you didn't spend on testing where the real revenue is.

Second, decision delay. A pricing test gets stuck behind a CTA test in the same queue. The team that should be moving fast on the bigger bet waits its turn.

Let me make the cost concrete. Say I've got a test sitting at 90% CTBC, and the team wants 95 before shipping. Closing that gap costs 3 more weeks of traffic.

If the page in question does $50,000 a week and the winning variant lifts conversion 3%, 3 weeks of waiting costs roughly $4,500, just to close a 5-point gap in confidence. That number's illustrative, not a client result.

Run your own traffic and lift numbers through a proper program ROI model, and the cost of over-waiting stops being theoretical. It becomes a figure you can put in front of whoever's asking why you shipped before 95%.

The goal is matching your rigour to what's actually on the line. Ask that before you touch the test setup: how important is this decision, really?

Rethinking Confidence by Decision Type

Ask how important the decision is, and 2 categories fall out immediately.

  1. Financial decisions, new pricing, margin changes, anything structural, deserve high confidence. Getting these wrong is expensive and hard to undo. 95% confidence earns its wait here.
  2. Marketing decisions, a headline, a CTA, a banner, sit further up the funnel. Their downstream impact on the business is limited. Chasing 95% confidence on these is overkill. A directional signal is enough to act on, and acting on it frees up the pipeline for the bets that deserve the full treatment.

At Speero, I lean on Bayesian analysis and a metric I call Confidence to Beat Control, CTBC. For programs without millions of monthly visitors, which is most of them, it balances time, effort, and impact better than frequentist methods manage.

Why Bayesian Beats Frequentist for Velocity

Frequentist versus Bayesian CTBC: two ways to read a test Frequentist: hands you a p-value, assumes a fixed sample size before you look, peeking early raises the false positive rate, a p-value of 0.06 is hard to act on. Bayesian CTBC: updates continuously as data comes in, no waiting period, hit your threshold and you stop, states a plain probability such as an 87 percent chance this variant beats control, which is easy to act on. WHY BAYESIAN BEATS FREQUENTIST Two Ways to Read a Test Frequentist confidence waits for a fixed sample. Bayesian updates as data comes in. FREQUENTIST (P-VALUE) Waits for a fixed sample size HOW IT BEHAVES → Hands you a p-value to interpret → Assumes a fixed sample before you look → Peek early and false positives climb A p-value of 0.06 is hard to act on. BAYESIAN (CTBC) Updates continuously as data arrives HOW IT BEHAVES → Updates continuously, no fixed sample → Hit your CTBC threshold, you stop → States a plain probability of winning "An 87% chance of beating control" is easy to act on. FIG. 01 · FREQUENTIST VS. BAYESIAN (CTBC) speero.com

Frequentist methods, the ones that hand you a p-value, have a problem if you're trying to move fast. It's called peaking. Check your results before you hit the sample size you planned for, and your false positive rate climbs. The method assumes you'll wait for a fixed number of sessions before you look. Almost nobody does. What you get instead is p-hacking, intentional or not.

Bayesian analysis sidesteps the whole issue. It updates continuously as data comes in, so you can watch the test without wrecking the integrity of the read. Hit your CTBC threshold, and you stop. No waiting period. No arbitrary sample size gatekeeping the moment your numbers are allowed to mean something.

It's also easier to explain. "There's an 87% probability this variant beats control" is something any product manager can act on. A p-value of 0.06 is not.

At Speero, I'll call a win above 90% CTBC on key tests, and flex depending on context. An 85% CTBC test might still get called flat while I flag the direction.

The threshold gets set before the test launches. Not after I've had a look at the numbers. Run the numbers yourself with a free significance calculator before you commit to one.

Pre-Registration: Agree the Rules Before You See the Data

Three habits pre-registration kills Pre-registering decision criteria before a test launches kills three habits: post-hoc justification, moving the goalposts once a test underperforms so the story still looks like learning; p-hacking, rerunning the analysis until something clears significance; and narrative spinning, calling a loss a strategic win because someone found an interesting segment. PRE-REGISTRATION Three Habits Pre-Registration Kills Write the rules before you see the data, not after. 01 Post-hoc justification Moving the goalposts once a test underperforms. 02 P-Hacking Rerunning the analysis until something clears significance. 03 Narrative spinning Calling a loss a strategic win because someone found an interesting segment. FIG. 02 · WHAT PRE-REGISTRATION PREVENTS speero.com

This is where most programs leave value sitting on the table. The discipline that actually matters happens one step earlier. Pre-registering your decision criteria before the test launches.

Write it down in advance. If CTBC clears 90%, ship it. Between 80 and 90, flag the direction and consider a follow-up.

Below 80, call it flat. The exact numbers shift by decision type. The principle doesn't: you write the rules before you see the data, not after.

Pre-registration kills 3 habits that quietly rot a program:

  • Post-hoc justification. Moving the goalposts after a test underperforms so the story still looks like learning.
  • P-hacking. Rerunning the analysis until something clears significance.
  • Narrative spinning. Calling a loss a strategic win because someone found an interesting segment.

It also forces something most programs avoid: taking losing tests as seriously as winning ones. A test that loses at 75% CTBC carries just as much information as one that wins at 92%.

Most of the real learning sits in the losses. Pre-registration is what makes sure you actually go get it.

The Decision Matrix: Matching Rigour to Risk

The decision matrix: matching rigour to risk Two axes set the confidence threshold. High reversibility and low impact or risk, such as copy, buttons, and image variants: fast-track at 80 to 85 percent CTBC, because being wrong costs little and waiting costs more. Low reversibility and high impact or risk, such as pricing, core flows, and the business model itself: hold out for 95 percent or better, because getting it wrong costs far more than the wait does. New features and big redesigns are mixed and follow whichever axis dominates. THE DECISION MATRIX Matching Rigour to Risk Reversibility and impact together set your confidence bar. Low High REVERSIBILITY Low High IMPACT & RISK HOLD FOR RIGOUR 95%+ CTBC Pricing, core flows, the business model itself Getting this wrong costs more than the wait does. FAST-TRACK 80–85% CTBC Copy, buttons, image variants Being wrong costs little. Waiting costs more. MIXED New features, big redesigns — follow whichever axis dominates. MIXED New features, big redesigns — follow whichever axis dominates. FIG. 03 · REVERSIBILITY × IMPACT → CONFIDENCE THRESHOLD speero.com

Before I launch anything, I need a way to categorise what the test is actually deciding. I map it across 2 axes in a documented decision matrix.

  1. Reversibility. If this turns out wrong, how hard is it to undo? A CTA change reverts in 10 minutes. A new pricing structure touches contracts and revenue recognition. Unwinding that takes months.
  2. Impact and risk. If this breaks the business model, how much damage does it do? A hero image swap barely registers. A change to the subscription model could be existential.

Put the two together and you get a working guide to confidence thresholds.

The Decision Matrix

Matching Rigour to Risk

Reversibility and impact together set your confidence bar.

Fig. 03 · Reversibility × impact → confidence threshold
Decision type Reversibility Impact & risk CTBC threshold Why
Fast-track High — reverts in minutes Low — copy, buttons, image variants 80–85% Being wrong costs little. Waiting costs more.
Hold for rigour Low — contracts, revenue recognition High — pricing, core flows, the business model 95%+ Getting this wrong costs far more than the wait does.
Mixed Depends on the specific change New features, big redesigns Follow the dominant axis A feature you can switch off fast leans toward speed. A redesign that changes how people see the brand leans toward rigour.

Swipe to see all columns →

  • High reversibility, low risk (copy, buttons, image variants): fast-track it. 80 to 85% CTBC is enough. Being wrong costs little. Waiting costs more.
  • Low reversibility, high risk (pricing, core flows, the business model itself): hold out for 95% or better. Getting this wrong costs far more than the wait does.
  • Mixed (new features, big redesigns): follow whichever axis dominates. A feature you can switch off fast leans toward speed. A redesign that changes how people see the brand leans toward rigour.

The matrix exists for one reason: to make your confidence bar match what's actually at stake.

What to Do With Flat Tests

Flat tests get treated like failures. They're the most consistently wasted output in this industry.

A well-powered flat test that pre-registered its hypothesis is telling you something. What it's telling you depends on how solid your pre-test thinking was.

  1. Grounded in real research, and a flat result probably means you found the right problem but the wrong fix. Go do focused research. Don't just run another version of the same idea.
  2. Grounded in a hunch, and a flat result might mean the problem you set out to solve doesn't actually exist. Go back upstream and find where the real friction lives.
  3. Underpowered, and a flat result tells you almost nothing. Check the traffic and the runtime before you draw any conclusion from it at all.

Every flat test earns a proper debrief and a clear call on whether to iterate or move on. Was the hypothesis grounded in evidence? Was the test powered enough to detect anything? What does the direction, however weak, point toward next?

Beyond A/B: The Velocity Toolbox

A/B testing proves or disproves a hypothesis better than anything else does. It's also the most expensive place in the program to learn something. Every session spent on a test that could've been killed earlier is a session you'll never get back.

Decision Velocity lives mostly in what happens before the A/B test launches.

Qualitative Research and Prototype Testing

Kill bad ideas before they reach the pipeline, and your win rate goes up. That's the fastest lever available. Qualitative research and prototype testing are how you pull it.

On anything business-critical, prototype testing catches usability problems that would otherwise wreck your A/B test results.

If a variant loses because the design confuses people, that loss is about execution, not your hypothesis. A moderated usability session before you ship costs a fraction of what a failed A/B test burns, and it means the test you eventually run is measuring an idea, not a UX mistake.

Pre/Post Testing

Pre/post compares before and after without a control. Sometimes you have no choice: low traffic, technical limits, a change you can't split cleanly. But it's a weak foundation for anything long-term.

Seasonality, campaigns, market shifts, all of it muddies the signal and hands you the wrong cause for the right effect. Use it when you have nothing else, and treat what it tells you as a hint, not a conclusion.

Multi-Arm Bandits

MABs shift traffic toward the better-performing variant automatically as the test runs. For optimisation work, content distribution, personalisation, engagement, they're genuinely good at what they do. They find winners fast.

The speed comes at a cost: learning rigour. MABs over-allocate to early leaders before the data has settled, which makes them a bad fit for anything business-critical. "Which of these 5 subject lines gets more opens" is a MAB question. "Should we restructure our pricing" is not, and a MAB will hand you a fast, confident answer to the wrong one. Use them to optimise. Leave the decisions that matter to something with more rigour behind it.

AI as a Velocity Multiplier

AI earns its place upstream, before the test ever launches. It raises the quality of what gets tested, which is exactly where most programs lose their time.

Narrow too early to one hypothesis or one design, and you shrink your learning surface. Lose, and you go back to the drawing board. Win, and you'll never know if a better version existed. Going all in too early is one of the quieter drags on velocity I see in program after program.

AI fixes this at the front end. Generate divergent ideas, designs, copy, creative directions, before you narrow anything down. Run 10 AI-generated headlines through a quick preference test and you'll cut the list to the strongest candidates in days, not weeks. Whatever survives carries a better shot at winning once it hits the A/B pipeline. Fewer failed tests. Fewer wasted sessions. Faster learning, full stop.

The workflow is simple. Generate a wide spread of directions with AI, across copy, design, positioning. Run a rapid preference test to cut the field.

Put your A/B traffic behind whatever survives, now carrying stronger priors than it would have otherwise.

AI's job here is shrinking the search space, so the A/B test stops carrying exploration that a cheaper, faster method should have handled already.

Learning Mode vs. Growth Mode

Most of the persistent velocity problems come from confusing 2 different jobs.

Learning mode builds knowledge. The goal is understanding: what's actually broken, why it's broken, what kind of fix might work. Hypothesis testing, qualitative research, prototypes, high-rigour A/B tests. Speed matters less than the quality of what you learn.

Growth mode captures value you've already proven works. The goal is exploitation: scale what's validated. MABs, optimisation tracks, fast iteration on known winners. Rigour can flex, because the hard question already got answered.

Mix the two and you pull the program in opposite directions at once. Push learning-mode work at growth-mode speed, and you scale ideas nobody's actually validated. Apply learning-mode rigour to every optimisation call, and you slow down without learning anything extra for the trouble.

Ask this before every test: am I learning something new, or extracting more value from something I already know works? The answer sets the tool, the threshold, and the timeline.

Execution Pitfalls That Kill Well-Powered Tests

A test can be correctly sized, correctly thresholded, and still tell you nothing if the execution is sloppy. A few killers show up again and again.

  1. Inconsistent variants change more than one thing at once. Shift the headline, the CTA, and the image together, and a win tells you the combination worked. A loss tells you nothing. Either way, you haven't learned what you set out to learn.
  2. Poor UX buries the effect you're trying to measure. A better value proposition wrapped in a confusing layout ends up testing the layout. The hypothesis never gets a fair hearing.
  3. Weak visual hierarchy means people never see the thing you're testing. The element isn't the most visually prominent thing on the page, so a chunk of your traffic walks straight past it. The test comes back flat, the hypothesis gets binned, and the real problem was that nobody saw the change in the first place.

Confirm a win actually holds up with a proper post-test validation method before you build on it. Quality beats quantity. Every time. 50 well-built tests a year will teach you more than 150 sloppy ones.

What This Means by Role

If you run growth or own a product, the tool selection is your call. Chasing 95% confidence on every micro-tweak is friction dressed up as rigour. Learn before you scale, and use the matrix to make the case for moving faster when stakeholders push back on a low-risk call.

If you manage the program, your job is pre-defining thresholds before launch, keeping learning experiments separate from optimisation tracks, and finding where velocity is actually bleeding out in the pipeline. The fix is almost never "run more tests."

If you're the strategist or the analyst, own the pre-test planning. That's where the real work happens: the hypothesis, the pre-registered criteria, the design review before results ever land. Make the rules explicit and public before the test launches, and p-hacking and post-hoc storytelling lose their cover.

If you design the experience, prototype the critical flows before they hit the pipeline. Every usability issue you catch in a moderated session is one less confound in the results. Talk to the analysts early, before you build, so what you're shipping is actually testable. Skip the scattergun redesign. Measurable change is meaningful change. Nothing else teaches you anything.

Where Decision Velocity Gets Built

Everything above assumes someone's actually sitting down to make the call. In practice, most of these decisions get made in two places: an ideation session deciding what to test next, and a quarterly review deciding what the last quarter's results actually mean.

Slow either one down and the statistics don't matter. The decision just arrives late. At Speero, I run AI-assisted systems for both.

MILES, our CRO ideation pipeline, takes a client's actual page, KPI, and swim lane and runs them through a phased ideation process, the same job good pre-test ideation is supposed to do, except it's pulling from real data instead of a whiteboard session.

A backlog scored with a PXL-style prioritization framework still decides what actually gets built. MILES just makes sure what lands in that backlog is worth scoring in the first place.

Q-BeRt, our automated QBR system, handles the other bottleneck. It reads a quarter's tests and research, maps each result back to the strategic pillar it was supposed to move, and assembles the review before anyone's had to reconstruct 3 months of decisions from memory.

The same evidence trail is what a program maturity audit is trying to establish. This one just runs continuously instead of once a year.

Neither tool replaces judgement. Programs that keep that evidence trail alive, instead of letting it sit in decks nobody reopens, make the next decision faster than programs that rebuild it from scratch every time. That's the point of building Decision Velocity into the system, instead of leaving it as a statistic you calculate after the fact.

The Principle Behind All of It

A program's value is how fast the team makes better decisions. Test count doesn't tell you that.

That means rigour where the risk demands it, speed where it doesn't, and a definition of success agreed before you launch. Every test. No exceptions.

Decision Velocity is the operating system underneath all of it: every decision gets exactly the scrutiny it's earned, and not a scrap more.

Pre-register your criteria. Match the threshold to the stakes. Kill the bad ideas early. Treat a flat test as data, never as a disappointment.

The programs that learn fastest are the ones that know which decisions are actually worth the wait.

Did you like this article?

(Your feedback helps us write better!) worst 1 - 10 best

Did the article resonate with you?

What aspects did you enjoy or find lacking?
Were there elements you felt we should've covered?
Thank you.
Oops! Something went wrong while submitting the form.

Related Posts

Who's currently reading The Experimental Revolution?