Experimentation / CRO

Experimentation and Testing Programs acknowledge that the future is uncertain. These programs focus on getting better data to product and marketing teams to make better decisions.

Research & Strategy

We believe that research is an integral part of experimentation. Our research projects aim to identify optimization opportunities by uncovering what really matters to your website users and customers.

Data and Analytics

90% of the analytics setups we’ve seen are critically flawed. Our data analytics audit services give you the confidence to make better decisions with data you can trust.

Speero Experimentation Blueprints

Experimentation Operating System (XOS) Blueprints help visualize organizational processes in order to optimize how a business delivers an experimentation program.

They have two parts: 1. they are decision support tools that are built on top of and customized, and 2. are connected with some program or business metric such as research velocity, or decision quality, speed, etc.

We present them as downloadable 'tools' (Figma, Miro, Decks, Docs, Sheets) for you to take, customize, and optimize your program with.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Filter blueprints by pillars of the XOS:

Why the default answer is no

An A/A test is usually asked for as a trust exercise. The problem is that the version people run, two identical experiences read on a conversion-rate p-value, is the weakest instrument available for the job. Three things go wrong with it.

A pass proves almost nothing. A non-significant A/A does not demonstrate the tool is sound, but only demonstrates the test lacked the power to find a difference. To rule out a difference as small as 0.1% relative on a 5% baseline conversion rate you would need roughly 298 million visitors per arm at 95% confidence and 80% power. AB Tasty put the same figure at around 300 million.

A failure is expected. At a 0.05 threshold, one A/A in twenty comes back significant while everything is working correctly. Teams that read that as a broken tool then distrust results they should have believed. Georgi Georgiev's objection is exactly this - the significance threshold already accounts for the outcome, so a rare significant result is a predictable event, not evidence of a fault.

Adding arms makes it worse. A/A/B and A/A/B/B designs are often proposed as a way to get validation and a real test in one shot. Every extra arm is another comparison and another chance of a false alarm, so the design that was meant to raise confidence lowers it.

But A/A tests do catch real bugs

The counter-evidence is strong enough that the tree keeps a live A/A on the table. Kohavi, Tangand Xu devote a full chapter to it with an empirical (not statistical) argument - the idea is useful because the tests fail so often in practice, and each failure forces a team to re-examine an assumption.

Ian Whitestone's write-up lists other catches, including a stats bug that doubled the false positive rate by skipping a multiple-comparisons adjustment, and redirect side effects where bots did not follow the redirect and created a sample ratio problem.

The resolution is not to argue about whether A/A tests work. It is to separate the question the A/Ais being asked to answer from the instrument that answers it cheapest.

Read an A/A on counts, not on conversions

If you do spend traffic on a live A/A, the sample ratio check on user counts is a far more sensitive bug detector than the conversion-rate comparison sitting next to it. Same visitors, same duration, very different resolving power.

Speero calculation. SRM detection uses a chi-square goodness-of-fit at alpha 0.001 with 80% power. The metric column is the minimum detectable relative difference on a 5% baseline conversion rate at alpha 0.05 with 80% power.
Users per armAllocation skew the SRM check catchesIn plain termsConversion difference the same test catches
5,00052.1 / 47.94.1% of users missing from one arm25.9%
10,00051.5 / 48.52.9% missing18.0%
30,00050.8 / 49.21.7% missing10.2%
100,00050.5 / 49.50.9% missing5.5%

At 30,000 per arm the count check flags a variant quietly losing 1.7% of its users. The conversion reading of the same data cannot see anything under about 10%. That gap is the whole reason for making SRM the primary read, and it is also why SRM belongs on every live A/B test, where it costs no extra traffic at all.

What actually justifies spending traffic on a live A/A

TriggerWhat justifies it
Trigger 01First implementationNew tool, or an existing tool on a site it has never run on. Nothing about the pipeline has been observed end to end yet.
Trigger 02MigrationClient-side to server-side, a tool swap, or a new personalisation or CDP layer sitting between the tool and the page.
Trigger 03Redirect or split-URL testingA documented SRM source. Bots and some users do not follow redirects, so one arm loses users before it is ever measured.
Trigger 04Numbers that never reconcileThe tool's counts and the analytics counts have never agreed and nobody knows why. That is a plumbing question, and an A/A isolates it.
Trigger 05Infrastructure changeNew CDN, new consent management platform, a rebuilt tag container, an SPA re-architecture. Any of these can silently break assignment.
Trigger 06SRM across multiple testsA pattern, rather than a single incident, points at the platform. Fabijan et al. call this out explicitly as a diagnostic rule of thumb.

How common are these failures:

  • Roughly 6% of experiments at Microsoft showed a sample ratio mismatch during the KDD study period. "A product running ten thousand experiments a year can expect to see at least one SRM per day."
  • The same paper catalogues 25 distinct root causes across five categories: assignment, execution, log processing, analysis, and interference. This is the taxonomy the diagnostic branch of the tree points to.
  • On the metric side, Microsoft found that the typical product had 10 to 15% of its metrics failing a p-value uniformity check under the null, and for some products the failure rate reached 30%. That is the case for the offline simulated A/A branch.

The three instruments, in cost order

CostInstrumentWhat it does
CheapestInstrumentation checkZero traffic, hours of work. Does the variant render in both arms, do the events fire, do exposures reconcile with analytics, does bucketing stay sticky across sessions and devices, is there flicker.
CheapOffline simulated A/AZero live traffic, uses history you already have. Resample 1,000+ times and confirm the p-value distribution is uniform. Catches skewed metrics, outlier sensitivity, rare events, and stats-engine bugs.
ExpensiveLive A/AReal traffic and real calendar time, and it can only be justified when the whole pipeline is unproven. Read it on counts and plumbing, and treat the conversion-rate p-value as a byproduct rather than the verdict.

Do you need an A/A test?

A single A/A test cannot prove your tool is sound. This decision tree shows when to run one, when an instrumentation check is enough, and when the sample ratio check already does the job.
View blueprint »
Artifact
Test & Learn

How do I craft a strong, testable hypothesis?

Strong experimentation programs are built on well-crafted hypotheses. The Hypothesis Creation Blueprint provides a structured, repeatable framework to help teams move from vague ideas to clear, testable statements. By focusing on the problem, defining a specific change, and identifying the expected outcome with rationale, teams reduce ambiguity and improve the strategic quality of their experiments. The framework also ensures that each test has a clear metric and business impact behind it. Over time, consistently using this blueprint supports a culture of evidence-based decision-making.

Use Cases:

- Helps teams move from vague ideas to clear, research-backed hypotheses that are easier to prioritize, scope, and validate.
- Ensures experiments are tied to a real user or business problem, improving test quality and increasing the likelihood of meaningful outcomes.
- Can be embedded into experimentation templates to standardize how hypotheses are written across the program.

Hypothesis creation blueprint

Strong experimentation programs are built on well-crafted hypotheses.
View blueprint »
Ritual
Planning & Process

What should my experimentation data stack look like considering the maturity of the program and the size of the company?

Every experimentation program is different. Some run a few experiments a month while others have hundreds of experiments running simultaneously. Some rely on a simple client-side tool while others have advanced server-side setups.

Different programs require different tools. This goes for the experimentation data stack, too.

While a small experimentation program can do everything using just their testing tool, more experienced users are looking for a richer dataset that's often found in their analytics tool (i.e. GA4). But that's nothing compared to the most advanced setups that are warehouse-native, data moves in multiple directions between various platforms and reporting is automated.

Use this blueprint to better understand what an ideal data stack might look like for your experimentation program.

Experimentation Data Stack Blueprint

Experimentation with different maturity level require different data stack and related tools.
View blueprint »
Artifact
Assessing and Scaling the Flywheel

What are the main Quality Assurance steps in A/B testing?

A step by step guide for Quality Assurance process in A/B testing.

How can you differentiate each part of QA?

What are the main areas and possible issues you should keep an eye on?

Use Cases:

- Quality Assurance is a must when you are A/B testing.
- While there are different tools and use cases, the main process needs to be the same for each experiments.
- Review and verify the setup, then test in different browsers and devices.

AB Testing QA process Blueprint

A step by step guide for Quality Assurance process in A/B testing.
View blueprint »
Ritual
Planning & Process

How do I present ROI and customer learnings across tests for a program?

The structure of the database of test learnings is important for communicating to stakeholders and assuring that decisions and actions are documented correctly. It can be a cultural tool above all, as it is a recipe for changing culture, from a data foundation. Not just arm waving about the theoretical benefits of democratized decision making.

Use Cases:

LOTS to take away from this structure below. Most of the takeaways I'd think should effect how you communicate to stakeholders.

For example:

1. A loss can equal a save

2. A flat test can still be 'implemented' (a win, it confirmed smth)

3. It shows the emphasis on RELATIVE effects across tests, which isn't talked enough in our industry dialog. Accuracy is over rated IMO, Precision FTW.

4. A financial model is critical. Create a goal tree map, then do a model to translate this into relative potential revenue. Use for BOTH prioritization and for reporting like this. It changes the game.

5. But don't only present $ numbers, also pair EVERY TIME with customer learning sentences, to TLDR what it meant for your customer's behavior and/or perceptions.

How to structure reporting across tests? Blueprint

The structure of the database of test learnings is important for communicating to stakeholders and assuring that decisions and actions are documented correctly.
View blueprint »
Artifact
Decision & Execution

How should I structure my experimentation program test reporting process?

Test reporting is critical to decision making, and also overall program velocity. The faster you can report out, the faster the decision can come. This is Agility as a metric. We think there are two sides to test reporting, the automated, and the bespoke or manual side. 1. BI tools like looker studio and even test tools themselves provide the automated side of things, and 2. the 'story telling' where different metrics are highlighted and the implications and insights are presented is the custom side of things. This blueprint acknowledges the need and balance for these to parts for experimentation test reporting.

Use Cases:

- increase awareness for what goes into a test reporting phase in the experimentation process

- align the team on who does what part, and what part is needs more work

- define your own programs component parts for this process

How should I structure my (automatic) test reporting? Blueprint

Test reporting is critical to decision making, and also overall program velocity. The faster you can report out, the faster the decision can come. This is Agility as a metric.
View blueprint »
Ritual
Decision & Execution

What are the differences between synchronous and asynchronous testing tool snippets?

What are the differences between synchronous and asynchronous testing tool snippets? Which one suits better for your website and testing program?

Use Cases:

While the difference between synchronous and asynchronous testing tool snippets may seem small, the actual impact this has on your website loading speeds, test flickering and overall user experience can be quite significant.

Synchronous vs Asynchronous Testing Tool Snippets - Pros and Cons Blueprint

What are the differences between synchronous and asynchronous testing tool snippets? Which one suits better for your website and testing program?
View blueprint »
Artifact
Planning & Process

Where are the key areas of opportunity on our website? How well does our website meet with key UX principles?

The definition of a “heuristic” is “a mental shortcut that allows people to solve problems and make judgments quickly and efficiently”. As such, our UX heuristic framework is made up of a set of guidelines which allow our team to assess and analyse any digital user experience and identify areas of opportunity for optimization.

Speero's UX heuristic framework was developed by combining and consolidating the frameworks used by the industry’s leading UX and CRO agencies. This resulted in a set of 60 guidelines across 5 heuristic themes;

Value: does the content communicate the value to the user?
‍Relevance: does the page meet user expectations in terms of content and design?
‍Clarity: is the content/offer on this page as clear as possible?
‍Friction: what is causing doubts, hesitations, uncertainties, and difficulties?
‍Motivation: does the content encourage and motivate users to take action towards the goal?

Frameworks similar to Speero's Heuristics Blueprint include:

MECLABS Conversion Index

Conversion's The Lever Framework

Conversion's The Lift Model

Use Cases:

- Assess any digital experience to understand and identify areas of opportunity for optimisation
- Use as a framework to tag and track action (JDIs, experiments, etc)

Heuristics Blueprint

The definition of a “heuristic” is “a mental shortcut that allows people to solve problems and make judgments quickly and efficiently”.
View blueprint »
Ritual
Assesment & Integration

How do I turn research insight into action?

Use this decision tree to help you to effectively categorize the insights generated via research. Effective categorization is where we turn insight into action and is the first step in developing an experimentation roadmap from research. Once you've categorized your insights, each list of insights can be dealt with accordingly, e.g. JDIs can be added to the development backlog or the next sprint, Instrument items can be handled by your analytics or development team, etc.

Use Cases:

- Turn research insight into an experimentation roadmap
- Create actionable workstreams for different teams
- Avoid the implementation crisis

Insight Categorisation Blueprint

Use this decision tree to help you to effectively categorize the insights generated via research.
View blueprint »
Ritual
Assesment & Integration

How do I align my program with business and customer goals?

Use Strategy Maps to orient your experimentation program. The focus will depend on the strategy and positioning of the brand.

Use Cases:

Understand if your experimentation is focused around Brand/Performance/Product marketing and Acquisition/Monetization/Rentention and whether your positioning is Sales/Marketing/Product led.

Strategy Maps Blueprint

Use Strategy Maps to orient your experimentation program. The focus will depend on the strategy and positioning of the brand.
View blueprint »
Artifact
Assesment & Integration

How to define different research objectives?

It can be challenging to effectively incorporate user research into experimentation programs on an ongoing basis. However, categorising all research initiatives into one of these three categories, can help you to plan research more effectively. Use this framework to communicate these three core research objectives.

Use Cases:

- To plan how to incorporate user research into your experimentation program on an ongoing basis
- To plan specific research initiatives, and keep them focused on a primary objective
- To communicate the different types of research being conducted to the wider business

Research Objectives Blueprint

It can be challenging to effectively incorporate user research into experimentation programs on an ongoing basis.
View blueprint »
Artifact
Assesment & Integration

What is the goal of the testing program?

There are different reasons to test and experiment, and they range from revenue to customer to process goals. CRO programs care about wins, they want to push for more money. CXO programs care about customers and thus are more focused on metrics around customer satisfaction and retention as (surrogates for money). XOS or 'experimentation' programs believe that a culture of using data and ultimately science is a better way to run an innovative company.

Use Cases:

- to help communicate with stakeholders where the focus should be

What is the goal of the testing program?

There are different reasons to test and experiment, and they range from revenue to customer to process goals.
View blueprint »
Artifact
Assessing and Scaling the Flywheel

Is my experimentation program profitable?

Assessing the revenue impact of your experimentation program is essential for informed decision making in business. Using a testing revenue model to measure ROI impact will help you understand the effectiveness of your strategies. Evaluating the direct revenue contribution of individual experiments allows you to determine where to allocate resources based on their effectiveness, ultimately enabling data-driven decision-making.

Use Cases:

- to measure the potential experimentation impact over time
- monitor the success of the experimentation program
- to acknowledge test performance decline over time

Testing revenue model

Assessing the revenue impact of your experimentation program is essential for informed decision-making in business.
View blueprint »
Ritual
Decision & Execution

Should I implement a data warehouse in my business or not?

Using Data Warehouse vs Not Using Data Warehouse Blueprint helps with just that. It shows that a data warehouse can help you bring all your data together to make better decisions and improve your business. Without a data warehouse, these tasks can be more challenging and less efficient.

Use Cases:

- The blueprints help you see why a data warehouse is good for your business.
- By looking at the table, you can decide if you need a data warehouse.
- The table guides you on what features to focus on if you're thinking about getting a data warehouse.

Using Data Warehouse vs Not Using Data Warehouse Blueprint

The "Using Data Warehouse vs Not Using Data Warehouse" blueprint compares using a data warehouse to not using one for your business.
View blueprint »
Artifact
Planning & Process

"Hypothesis Testing" and "Do No Harm" treatments represent different experimental goals, hence it is important to differentiate them to ensure correct statistical analysis and interpretation of results, as well as increase your testing velocity. Hypothesis testing is used to test whether one variation is significantly better than another, while a "Do No Harm" treatment is used to test whether one variant is not significantly worse than another by a predetermined margin.

Use Cases:

- Estimating the duration of an experiment
- Balancing a portfolio of experiments in your program
- Increasing your testing velocity

Hypothesis Testing vs “Do No Harm” Treatments in A/B Testing

"Hypothesis Testing" and "Do No Harm" treatments represent different experimental goals, hence it is important to differentiate them to ensure correct statistical analysis and interpretation of results.
View blueprint »
Artifact
Test & Learn

Based on my website's traffic and conversion numbers, how many experiments can I run at each step of the user journey?

A/B testing is interconnected with statistics and statistics always requires a certain sample size to draw any sort of meaningful results. Using a test bandwidth calculator and the —Where and how should I test to make the most money? Blueprint—you can determine whether your website's traffic is fit for A/B testing and whether you have the ability to go deeper into segments and down-the-funnel metrics.

Use Cases:

- Is your website eligible for A/B testing based on its current traffic and conversion volumes?
- Is it feasible for you to run experiments with smaller impact or should you focus on high-impact experiments higher in the funnel?

Where and How Should I Test to Make the Most Money? Blueprint

Based on my website's traffic and conversion numbers, how many experiments can I run at each step of the user journey?
View blueprint »
Artifact
Assesment & Integration