Experimentation / CRO

Experimentation and Testing Programs acknowledge that the future is uncertain. These programs focus on getting better data to product and marketing teams to make better decisions.

Research & Strategy

We believe that research is an integral part of experimentation. Our research projects aim to identify optimization opportunities by uncovering what really matters to your website users and customers.

Data and Analytics

90% of the analytics setups we’ve seen are critically flawed. Our data analytics audit services give you the confidence to make better decisions with data you can trust.

The Experimentation Center of Excellence, Deconstructed

Illustration for 'Built to Make Itself Obsolete' featuring Ben Labay, CEO of Speero, alongside a donut chart showing RS Group tested only 5% of shipped work before building a Center of Excellence.
The Diagnosis: An experimentation Center of Excellence doesn't fail because a company centralizes testing. It fails because nobody plans for what happens after centralization works, so a structure built to fix one bottleneck stays in place long after it's solved it.
RS Group's centralized team was testing 5% of what shipped while 95% went out untested. Beerwulf's UX team was producing reports other departments weren't acting on. Different companies, same pattern: a structure with no exit plan.
The fix across all ten interviews isn't a different reporting line. It's treating the CoE as a bridge with a known destination, and building the incentives that make letting go of it possible.

Ask ten experimentation leaders how to structure a testing program and you'll get ten different org charts.

Ask them what an experimentation center of excellence is actually for, and something strange happens. The answers converge.

Stewart Ehoff, commercial growth director at RS Group, a FTSE 100 engineering firm, describes it as governance, not generation.

Luis Trindade, principal product manager at Farfetch, a fashion marketplace with 6,000 employees, calls it teaching people to fish instead of handing them fish.

Liam Furnam, a data scientist who worked inside both Meta and Booking.com, describes it as one stop on a longer road toward full team ownership.

And Marty Cagan, the product management writer known for calling out bad organizational design, argues that most centers of excellence get the balance wrong in one direction or the other.

Ten Interviews, Ten Paths

How Ten Companies Built (and Evolved) Their CoE

Every interview traces a similar arc: centralize to fix a real problem, then hand ownership back to product teams.

Company · starting problem · CoE model · what replaced it
Company Starting problem CoE model What replaced it
RS Group A centralized team tested only 5% of what shipped; the other 95% went out disconnected from product strategy. The center stopped building tests and shifted to governing tools, process, and a shared test database across 12+ squads in 31 countries. Product teams now own ideation and execution; the center still runs test analysis until squads have dedicated analysts of their own.
Specsavers A three-person team was stretched across markets and products; the real bottleneck was ideation, not execution. A standardized four-phase framework, a briefing template for new ideas, a RASCI/governance model, and OKRs split between capability-building and test velocity. Templates and governance distributed the process itself, instead of growing headcount to keep up.
Online Dialogue Not a single company's case — Ruben de Boer consults across many clients, each starting from a different place. A spectrum of models depending on the client: full-service (the CoE designs and runs tests), coaching-only (teaches teams to self-serve), or a "traveling circus" building skills team by team. No fixed replacement — the model is matched to each client's readiness rather than standardized.
Constant Contact At a prior, unnamed employer, product teams came to the CoE with widely varying levels of need. A coaching CoE that set up tools, guided reporting, and built learning roadmaps for teams. Some specialists were embedded directly into product pods for deeper support — it helped those pods, but left fewer people to support the rest of the org.
Farfetch Starting the CoE journey seven years ago at 6,000+ employees, marketing and engineering wanted different tools with no shared approach. A small central team trains an "Experimentation Champions Network" of PMs, analysts, and designers instead of centralizing execution. Tool fragmentation gave way to shared standards; execution stayed distributed through the champions network.
Diligent (ex-Uber) Traditional organizations route ideas through JIRA-style workflows without validating them first, so testing stays siloed. The CoE acts as a bridge, educating teams and standardizing tools to move the org toward a "product operating model." At Uber, cited as the model's end state, the bridge became unnecessary — "100% of things" run through A/B testing inside autonomous pods.
Vista Testing ran top-down and project-managed, before teams had the maturity to own it themselves. The center runs no tests itself; it supports and enables teams that own their own metrics and problem spaces. As teams matured, the center's role shifted from running tests to strategy, coaching, and resourcing.
ING Not stated as a problem — cited as a model that worked well at that stage. A fully centralized team of dedicated specialists, running "over a couple of hundred tests per year." Not stated — presented as still effective for a smaller, earlier-stage org, not something later replaced.
Booking.com Segment-based teams (hotel owners, travelers, service) produced underpowered tests; expanding into flights and cars made a team-per-product-line model unsustainable. All the separate teams were centralized into one experimentation department for scale. Centralization risked an "ivory tower" disconnect, partly addressed with embedded "ambassador" roles; a later quality-metrics push proved harder to sustain than incentive-based approaches.
Meta Not framed as a "before" problem — presented as the working model. Teams had to run a backtest (launch to 100%, hold back a control group) to claim any revenue impact; a central team managed shared "test slots" across parallel experiments. Not stated — described as Meta's steady-state practice, not a structure that itself got replaced.

Swipe to see all columns →

Note: the Constant Contact and Online Dialogue rows reflect what those interviewees described in their own words — a prior, unnamed employer's case in Rommil Santiago's interview, and a consultant's general view across clients in Ruben de Boer's — rather than a documented case at Constant Contact or Online Dialogue themselves.

Speero spent ten interviews talking to people who built these programs.

The companies ranged from a 40-year-old optics retailer to Uber, Farfetch, and a Heineken subsidiary that makes home beer taps.

None of them set out to describe the same thing.

But a pattern showed up anyway, centered on what a center of excellence is actually built to do.

For most programs, that's the opposite of what they assume going in.

An Experimentation Center of Excellence With a Built-In Expiration Date

An experimentation center of excellence works like a bridge across an organization's maturity curve.

Every single interviewee described it that way, in their own words.

The Center of Excellence lifecycle A three-stage process diagram showing how experimentation programs move from decentralized testing, through a centralized Center of Excellence, to distributed empowered teams. Where a Center of Excellence Fits in the Lifecycle Decentralized Ad hoc testing, no shared standards, inconsistent results. Centralized / CoE One team sets standards, governs or runs tests, fixes the inconsistency. Distributed / Empowered Teams Guilds, champions, or empowered pods own testing. CoE shifts to coaching. speero.com

Stewart Ehoff, commercial growth director at RS Group, a FTSE 100 engineering solutions provider, put it plainly. The CoE's job is to take a squad "from zero to one experiment," then loosen its grip as that squad matures.

Vista's Kevin Anderson, senior product manager of experimentation, described the CoE as a transition device. It moves a company from top-down, project-managed testing toward an empowered product culture. After that shift, the central team's role becomes strategy, hiring, and coaching instead of running tests.

Diligent's Dan Layfield, director of product management who spent years at Uber, called the CoE a transition phase toward making experimentation universal across every product team.

Farfetch's Luis Trindade, principal PM of experimentation, summed up the philosophy in five words: teach people how to fish. Even the interview that pushed back hardest on centralization agreed with the premise.

Tim Thijsse, a senior CXO specialist who ran the UX function at Beerwulf, didn't just loosen his centralized team's grip. He dissolved it entirely. In its place came an informal "Guild" of people who shared an interest in experimentation and customer insight, with every individual contributor moved fully into a product team.

If you're building a program right now and treating the CoE as the finish line, that's the first thing worth unlearning.

The finish line is a company where every team runs rigorous experiments without a central group doing it for them.

Why Centralization Comes First (and Why It Should)

None of this means centralization is a mistake.

Every company in this series centralized for a real reason. Usually because decentralized testing was quietly failing in a way leadership hadn't noticed yet.

At RS Group, the fully decentralized state meant only 5% of shipped products got tested at all. The other 95% bypassed experimentation entirely and disconnected from any broader strategy.

RS Group's test coverage gap before a Center of Excellence Donut chart showing that 5% of RS Group's shipped products were tested and 95% shipped without any testing, before RS Group built a Center of Excellence. RS Group's Test Coverage Before a Center of Excellence 95% untested Tested: 5% Source: Speero interview with RS Group's commercial growth director, before Center of Excellence rollout. speero.com

At Booking.com, Liam Furnam, a data scientist who later worked at Meta and now works at RealFi, described decentralized teams built around customer segments: hotel owners, individual travelers, service teams. That structure ran into a variance problem

Large hotel chains behave nothing like individual hosts, so tests kept coming back underpowered. When Booking.com expanded into flights and car rentals, running separate teams for every segment stopped being sustainable. Centralization improved scale, even though it risked losing touch with individual product needs.

At Beerwulf, before the Guild model existed, the centralized UX team was producing insight and experiment reports that other departments simply weren't acting on. Its own people were stretched thin across multiple product teams and a separate UX board at the same time.

Specsavers' Melanie Kyrklund, global head of experimentation who previously built optimization programs at Booking.com, Staples Europe, and Liberty Global, faced a different version of the same problem. Her three-person team's real bottleneck was ideation: too many untested ideas, not enough structure to evaluate them before they entered the pipeline.

And at ING, a fully centralized model run by a single dedicated team supported "over a couple of hundred tests per year." That's proof centralization can move fast when a company is small enough for one team to see the whole picture.

The lesson is more specific than "centralize" or "don't centralize."

Centralization solves a nameable problem: inconsistent standards, underpowered tests, disconnected ideation, or no process at all.

Where Centralization Starts to Cost You

The trouble starts once the company outgrows the problem centralization was built to solve.

Constant Contact's Rommil Santiago, senior director of product experimentation and founder of Experiment Nation, named the trade-off directly. A CoE works well when teams need coaching on execution basics. But scaling forces an ugly choice: stay centralized and create bottlenecks and knowledge silos, or embed people in pods and lose the ability to serve the wider organization.

"The structure ultimately depends on the company." (Rommil Santiago, Senior Director of Product Experimentation, Constant Contact)

Marty Cagan, a partner at Silicon Valley Product Group, frames the same tension at the extremes. Fully centralized CoEs prevent teams from learning and slow execution down. Fully decentralized setups create inefficiency and duplicated effort.

His "happy medium" is an expertise-based CoE that coaches and guides rather than running the tests itself. He argues that model holds up better under budget pressure than a roster of embedded specialists, who tend to get cut the moment the budget tightens.

Cagan goes further than most interviewees. He warns that even a well-run CoE can quietly cap a company's ambition if it treats experimentation as pure optimization. He calls this a "bug, not a feature" of how most organizations think about testing.

Real product management, in his view, requires both low-risk optimization and higher-risk innovation work. Amazon Prime is his example of the difference. The breakthrough there came from deep, high-risk exploration of shipping logistics and cost structures, years before any button got tested:

  1. Optimization is value capture.
  2. Innovation is value creation.
  3. A CoE that only ever runs the first kind of test is doing useful, limited work.

Cagan also takes aim at a popular framework for the same reason. The Double Diamond implies teams should spend roughly equal time on problem discovery and solution development. At scale, he argues, that creates organizational chaos. He prefers a pyramid: leadership identifies the critical problems, and teams discover the solutions.

Why Incentives Decide Whether the Structure Works

The org chart matters less than what people are measured on. That idea runs underneath every one of these ten interviews.

Santiago put it plainly. Teams measured on outcome metrics like retention and revenue test aggressively. Teams measured on delivery and completion tend to stop testing the moment something ships.

Furnam described how Meta enforced this at the highest level. Teams couldn't claim credit toward quarterly revenue goals without running a backtest first: launch the feature fully, hold it back for 5% of users, and measure the real difference.

"If you wanted to claim impact from your launch, you had to run an experiment." Revenue goals were tied directly to those backtest estimates. Teams knew they had one shot to prove impact, and that incentive produced real rigor instead of hopeful guessing. Meta also managed shared "test slots," so individual teams couldn't run experiments freely and risk interaction effects that would corrupt everyone else's data.

Booking.com tried a different lever: quality metrics tracking whether teams and departments were following proper experimental practice. It's a reasonable idea in theory. In practice, it proved harder to enforce and more prone to loopholes than tying incentives directly to revenue.

Vista's approach to the same problem was structural rather than punitive. Before anything else, the company spent six months mapping metrics from top-level KPIs all the way down to individual team metrics. Anderson described that foundation as essential before teams could set meaningful quarterly OKRs.

"The insights and learnings should come from the product teams, and bubble upwards toward the product leaders, which informs the future offers." Kevin Anderson.

Farfetch built something similar under the name "smart KPIs." Each autonomous business domain blends business and consumer metrics, and in-domain analysts translate test results into financial impact. A single homepage experiment can be traced all the way to a company objective.

Specsavers split its OKRs into two deliberate buckets. One tracks capability-building: training, competence, maturity. The other tracks experiment velocity and quality. That structure let leadership see commercial value even while teams were still climbing the learning curve.

These are incentive decisions, and they show up in every interview that talks about why a program actually changed behavior instead of just changing reporting lines.

What Actually Replaces the Experimentation Center of Excellence

So what do you build once centralization has done its job and you're ready to loosen the grip?

The interviews describe roughly the same three tools, wearing different names.

Champions and guilds. Farfetch built an "Experimentation Champions Network" of product managers, analysts, and designers spread across product, platform services, and partner companies. The central team trains and engages them regularly, so expertise sits closer to the business without losing consistency.

Beerwulf went a step further and dissolved its formal team altogether. Its Guild ran on weekly stand-ups, three-week OKR review cycles focused on process quality rather than output volume, and a shared backlog despite people being spread across different teams.

The Guild's real output was frameworks, templates, and guidelines that product teams could run with on their own. As Thijsse noted, this only works "if the company's vision and strategy are already aligned with a customer-centric focus." At Beerwulf, that meant customer satisfaction was a company-wide KPI, not just a UX team metric.

Governance instead of generation. RS Group's version of this was explicit. The CoE's job is setting the standard that "everyone must use the same tools, technologies, and processes," leaving test ideas and experiment builds to the product teams themselves.

Every experiment feeds into one central database for visibility. Onboarding follows a "paint-by-numbers" process covering what experimentation means, its business value, procedures, KPIs, and who's responsible for what.

The team splits roles between specialists (communication, stakeholder management, product experience) and data specialists (quality, validation, analysis, guardrails), rather than hunting for unicorns who can do both.

Documented process over tribal knowledge. Specsavers built a briefing template that captures how a problem was identified, what the problem actually is, and whether it's been quantified.

The goal was stopping weak ideas from entering the pipeline. The company standardized on Clickvalue's four-phase framework (problem discovery, problem validation, solution discovery, solution validation) as shared language for training and alignment. A RACI-style matrix backed it up, clarifying who's responsible, accountable, supportive, and consulted at each stage.

Online Dialogue's Ruben de Boer, lead experimentation consultant, described the range of models this can take. Full-service, where the CoE handles design, analysis, and hypothesis generation directly.

Coaching-focused, where workshops teach product teams to run their own tests. And a "traveling circus" model that moves team by team building velocity skills before moving on.

"You aren't experimenting to build a CoE." Ruben de Boer.

The center of excellence is a tool for improving your tools, teams, process, and culture. It was never supposed to be the goal itself.

Where This Leaves Your Program

Ten companies, ten different starting points, and the same three-part story. Decentralized testing breaks down. Centralization fixes it and creates new problems of its own.

The fix for those new problems is teaching, standardizing, and getting out of the way.

Layfield, describing how Uber ran things, put a number on what full commitment looks like. "100% of things" got A/B tested there, from backend infrastructure updates to front-end changes. Experimentation was woven into daily operations, not treated as a special project waiting on a JIRA ticket.

That's the state every interview in this series is pointing toward. Some companies call it a Guild. Others call it a Champions Network, or an empowered product team, or nothing at all because it's just how the company works.

If your program is still running everything through a central team, that's likely the right structure for where you are right now.

The work is making sure your experimentation center of excellence doesn't stay that way past the point where it's still solving a real problem.

Read the full series for the details behind every one of these companies. Start with how RS Group rebuilt testing across 31 countries, or go straight to Marty Cagan's case for why optimization alone isn't enough.

Did you like this article?

(Your feedback helps us write better!) worst 1 - 10 best

Did the article resonate with you?

What aspects did you enjoy or find lacking?
Were there elements you felt we should've covered?
Thank you.
Oops! Something went wrong while submitting the form.

Related Posts

Who's currently reading The Experimental Revolution?