The Diagnosis: An experimentation Center of Excellence doesn't fail because a company centralizes testing. It fails because nobody plans for what happens after centralization works, so a structure built to fix one bottleneck stays in place long after it's solved it.
RS Group's centralized team was testing 5% of what shipped while 95% went out untested. Beerwulf's UX team was producing reports other departments weren't acting on. Different companies, same pattern: a structure with no exit plan.
The fix across all ten interviews isn't a different reporting line. It's treating the CoE as a bridge with a known destination, and building the incentives that make letting go of it possible.
Ask ten experimentation leaders how to structure a testing program and you'll get ten different org charts.
Ask them what an experimentation center of excellence is actually for, and something strange happens. The answers converge.
Stewart Ehoff, commercial growth director at RS Group, a FTSE 100 engineering firm, describes it as governance, not generation.
Luis Trindade, principal product manager at Farfetch, a fashion marketplace with 6,000 employees, calls it teaching people to fish instead of handing them fish.
Liam Furnam, a data scientist who worked inside both Meta and Booking.com, describes it as one stop on a longer road toward full team ownership.
And Marty Cagan, the product management writer known for calling out bad organizational design, argues that most centers of excellence get the balance wrong in one direction or the other.
Speero spent ten interviews talking to people who built these programs.
The companies ranged from a 40-year-old optics retailer to Uber, Farfetch, and a Heineken subsidiary that makes home beer taps.
None of them set out to describe the same thing.
But a pattern showed up anyway, centered on what a center of excellence is actually built to do.
For most programs, that's the opposite of what they assume going in.
An Experimentation Center of Excellence With a Built-In Expiration Date
An experimentation center of excellence works like a bridge across an organization's maturity curve.
Every single interviewee described it that way, in their own words.
Stewart Ehoff, commercial growth director at RS Group, a FTSE 100 engineering solutions provider, put it plainly. The CoE's job is to take a squad "from zero to one experiment," then loosen its grip as that squad matures.
Vista's Kevin Anderson, senior product manager of experimentation, described the CoE as a transition device. It moves a company from top-down, project-managed testing toward an empowered product culture. After that shift, the central team's role becomes strategy, hiring, and coaching instead of running tests.
Diligent's Dan Layfield, director of product management who spent years at Uber, called the CoE a transition phase toward making experimentation universal across every product team.
Farfetch's Luis Trindade, principal PM of experimentation, summed up the philosophy in five words: teach people how to fish. Even the interview that pushed back hardest on centralization agreed with the premise.
Tim Thijsse, a senior CXO specialist who ran the UX function at Beerwulf, didn't just loosen his centralized team's grip. He dissolved it entirely. In its place came an informal "Guild" of people who shared an interest in experimentation and customer insight, with every individual contributor moved fully into a product team.
If you're building a program right now and treating the CoE as the finish line, that's the first thing worth unlearning.
The finish line is a company where every team runs rigorous experiments without a central group doing it for them.
Why Centralization Comes First (and Why It Should)
None of this means centralization is a mistake.
Every company in this series centralized for a real reason. Usually because decentralized testing was quietly failing in a way leadership hadn't noticed yet.
At RS Group, the fully decentralized state meant only 5% of shipped products got tested at all. The other 95% bypassed experimentation entirely and disconnected from any broader strategy.
At Booking.com, Liam Furnam, a data scientist who later worked at Meta and now works at RealFi, described decentralized teams built around customer segments: hotel owners, individual travelers, service teams. That structure ran into a variance problem
Large hotel chains behave nothing like individual hosts, so tests kept coming back underpowered. When Booking.com expanded into flights and car rentals, running separate teams for every segment stopped being sustainable. Centralization improved scale, even though it risked losing touch with individual product needs.
At Beerwulf, before the Guild model existed, the centralized UX team was producing insight and experiment reports that other departments simply weren't acting on. Its own people were stretched thin across multiple product teams and a separate UX board at the same time.
Specsavers' Melanie Kyrklund, global head of experimentation who previously built optimization programs at Booking.com, Staples Europe, and Liberty Global, faced a different version of the same problem. Her three-person team's real bottleneck was ideation: too many untested ideas, not enough structure to evaluate them before they entered the pipeline.
And at ING, a fully centralized model run by a single dedicated team supported "over a couple of hundred tests per year." That's proof centralization can move fast when a company is small enough for one team to see the whole picture.
The lesson is more specific than "centralize" or "don't centralize."
Centralization solves a nameable problem: inconsistent standards, underpowered tests, disconnected ideation, or no process at all.
Where Centralization Starts to Cost You
The trouble starts once the company outgrows the problem centralization was built to solve.
Constant Contact's Rommil Santiago, senior director of product experimentation and founder of Experiment Nation, named the trade-off directly. A CoE works well when teams need coaching on execution basics. But scaling forces an ugly choice: stay centralized and create bottlenecks and knowledge silos, or embed people in pods and lose the ability to serve the wider organization.
"The structure ultimately depends on the company." (Rommil Santiago, Senior Director of Product Experimentation, Constant Contact)
Marty Cagan, a partner at Silicon Valley Product Group, frames the same tension at the extremes. Fully centralized CoEs prevent teams from learning and slow execution down. Fully decentralized setups create inefficiency and duplicated effort.
His "happy medium" is an expertise-based CoE that coaches and guides rather than running the tests itself. He argues that model holds up better under budget pressure than a roster of embedded specialists, who tend to get cut the moment the budget tightens.
Cagan goes further than most interviewees. He warns that even a well-run CoE can quietly cap a company's ambition if it treats experimentation as pure optimization. He calls this a "bug, not a feature" of how most organizations think about testing.
Real product management, in his view, requires both low-risk optimization and higher-risk innovation work. Amazon Prime is his example of the difference. The breakthrough there came from deep, high-risk exploration of shipping logistics and cost structures, years before any button got tested:
- Optimization is value capture.
- Innovation is value creation.
- A CoE that only ever runs the first kind of test is doing useful, limited work.
Cagan also takes aim at a popular framework for the same reason. The Double Diamond implies teams should spend roughly equal time on problem discovery and solution development. At scale, he argues, that creates organizational chaos. He prefers a pyramid: leadership identifies the critical problems, and teams discover the solutions.
Why Incentives Decide Whether the Structure Works
The org chart matters less than what people are measured on. That idea runs underneath every one of these ten interviews.
Santiago put it plainly. Teams measured on outcome metrics like retention and revenue test aggressively. Teams measured on delivery and completion tend to stop testing the moment something ships.
Furnam described how Meta enforced this at the highest level. Teams couldn't claim credit toward quarterly revenue goals without running a backtest first: launch the feature fully, hold it back for 5% of users, and measure the real difference.
"If you wanted to claim impact from your launch, you had to run an experiment." Revenue goals were tied directly to those backtest estimates. Teams knew they had one shot to prove impact, and that incentive produced real rigor instead of hopeful guessing. Meta also managed shared "test slots," so individual teams couldn't run experiments freely and risk interaction effects that would corrupt everyone else's data.
Booking.com tried a different lever: quality metrics tracking whether teams and departments were following proper experimental practice. It's a reasonable idea in theory. In practice, it proved harder to enforce and more prone to loopholes than tying incentives directly to revenue.
Vista's approach to the same problem was structural rather than punitive. Before anything else, the company spent six months mapping metrics from top-level KPIs all the way down to individual team metrics. Anderson described that foundation as essential before teams could set meaningful quarterly OKRs.
"The insights and learnings should come from the product teams, and bubble upwards toward the product leaders, which informs the future offers." Kevin Anderson.
Farfetch built something similar under the name "smart KPIs." Each autonomous business domain blends business and consumer metrics, and in-domain analysts translate test results into financial impact. A single homepage experiment can be traced all the way to a company objective.
Specsavers split its OKRs into two deliberate buckets. One tracks capability-building: training, competence, maturity. The other tracks experiment velocity and quality. That structure let leadership see commercial value even while teams were still climbing the learning curve.
These are incentive decisions, and they show up in every interview that talks about why a program actually changed behavior instead of just changing reporting lines.
What Actually Replaces the Experimentation Center of Excellence
So what do you build once centralization has done its job and you're ready to loosen the grip?
The interviews describe roughly the same three tools, wearing different names.
Champions and guilds. Farfetch built an "Experimentation Champions Network" of product managers, analysts, and designers spread across product, platform services, and partner companies. The central team trains and engages them regularly, so expertise sits closer to the business without losing consistency.
Beerwulf went a step further and dissolved its formal team altogether. Its Guild ran on weekly stand-ups, three-week OKR review cycles focused on process quality rather than output volume, and a shared backlog despite people being spread across different teams.
The Guild's real output was frameworks, templates, and guidelines that product teams could run with on their own. As Thijsse noted, this only works "if the company's vision and strategy are already aligned with a customer-centric focus." At Beerwulf, that meant customer satisfaction was a company-wide KPI, not just a UX team metric.
Governance instead of generation. RS Group's version of this was explicit. The CoE's job is setting the standard that "everyone must use the same tools, technologies, and processes," leaving test ideas and experiment builds to the product teams themselves.
Every experiment feeds into one central database for visibility. Onboarding follows a "paint-by-numbers" process covering what experimentation means, its business value, procedures, KPIs, and who's responsible for what.
The team splits roles between specialists (communication, stakeholder management, product experience) and data specialists (quality, validation, analysis, guardrails), rather than hunting for unicorns who can do both.
Documented process over tribal knowledge. Specsavers built a briefing template that captures how a problem was identified, what the problem actually is, and whether it's been quantified.
The goal was stopping weak ideas from entering the pipeline. The company standardized on Clickvalue's four-phase framework (problem discovery, problem validation, solution discovery, solution validation) as shared language for training and alignment. A RACI-style matrix backed it up, clarifying who's responsible, accountable, supportive, and consulted at each stage.
Online Dialogue's Ruben de Boer, lead experimentation consultant, described the range of models this can take. Full-service, where the CoE handles design, analysis, and hypothesis generation directly.
Coaching-focused, where workshops teach product teams to run their own tests. And a "traveling circus" model that moves team by team building velocity skills before moving on.
"You aren't experimenting to build a CoE." Ruben de Boer.
The center of excellence is a tool for improving your tools, teams, process, and culture. It was never supposed to be the goal itself.
The Two Axes Your Experimentation Center of Excellence Sits On
Ten leaders, ten org charts, and none of them contradict each other.
They only look contradictory because the argument is usually held on one axis. Centralized on the left, decentralized on the right, every program supposedly crawling from one end toward the other.
Speero's work with Nils Stotz and Paul Drews at Leuphana University, and with Lukas Vermeer of Vista, drew the second axis that makes ten answers reconcile into one picture. A taxonomy for structuring experimentation teams plots team structure against operating model. The vertical axis asks whether your company runs a product operating model, where teams own outcomes, or a feature-management model, where teams deliver a specified scope by a date.
Four positions come out of that, and they behave differently enough to earn names. The org charts blueprint covers the three shapes an experimentation team can take. The quadrants tell you which of those shapes your operating model can actually hold up.
Product-driven centralization. One rigorous team serving product-minded pods. Experimentation gets used to shape direction rather than to sign off decisions already made. It also bottlenecks by design, because a single team ends up consulting for more stakeholders than it can serve. ING's dedicated team running over a couple of hundred tests a year sits here and works fine, because the company is small enough for one team to see the whole picture.
Product-driven decentralization. Execution lives inside cross-functional teams, running on a shared platform and shared metric definitions. This is the state everyone in this series is describing when they talk about handing over, and it's the only quadrant where test volume and decision quality climb together.
Feature-led centralization. A central team polices method inside a company that ships to a plan. Tests get scoped to launches and milestones, and features go out regardless of what the test said. Most centers of excellence end up here, which is where the bureaucratic auditor reputation was earned.
Feature-led decentralization. Capacity is spread everywhere and nothing holds it together. Methods diverge, tooling fragments, and results stop being poolable across teams. RS Group's pre-CoE state is the recognizable version, with 95% of shipped product going out untested and disconnected from any strategy.
An experimentation center of excellence sits across the middle of this map rather than occupying a fifth box on it. It carries tooling, definitions, and training from one quadrant toward the next. That is the load-bearing version of the bridge argument, because a bridge needs two banks and a direction of travel.
Reading the map that way explains the failure mode more precisely than saying a team calcified.
Most centers of excellence get chartered to move a company from product-driven centralization to product-driven decentralization. Growth pressure arrives, the company imports feature-management governance to keep delivery predictable, and the program slides down into feature-led centralization instead. The same team built to distribute capability is now the review board, and climbing back out takes an operating model change and a structural change at the same time.
Booking.com is your reason to stop treating the horizontal axis as a ratchet. Its segment-based teams were already decentralized and producing underpowered tests, so it centralized into a single experimentation department as it expanded into flights and car rentals. Then it embedded ambassadors to buy back the product closeness that centralization cost. Moving left was the right call.
Champions networks, guilds, and ambassador programs are one mechanism wearing three names, and that mechanism is what makes the right-hand column survivable at all. Beerwulf's Guild held because customer satisfaction was already a company-wide KPI rather than a UX team metric. Decentralize with no connective mechanism and you land in the wild west whatever the intent was.
Two fast diagnostics will place you on the map before you argue about structure.
The first is the vertical axis, and it has money attached to it. McKinsey's Operating Model Index research scored more than 400 public companies and found the top quartile showed 60% higher shareholder returns and 16% higher operating margins than the bottom half. Funding tied to measurable goals showed up among the largest gaps between the top and bottom performers, which is the operating model deciding what a test is allowed to change.
The second is culture, and you can run it as one question. What happened to the last unexpected result? Westrum's organizational culture model sorts the answers into three. Pathological companies hoard information or distort it for political reasons. In a bureaucratic one, departments defend their turf and insist on doing things by the book. Generative companies subordinate everything to performance. Culture predicts how information moves, which is why one identical test report lands three different ways in three companies with matching org charts.
Cagan's warning about optimization capping ambition is a third dimension on the same map. A program can sit in the right quadrant and still spend every test on tweaks.
Pointing your experimentation program metrics at outcomes instead of output is usually the first visible move. A goal tree map linking team metrics to business goals is how most programs find out their measurement stops two levels above the team doing the work.
Placing yourself on this map tells you where you are. What has to change before you can hand over is a separate question, and it has four specific answers.
Four Gates That Tell You the Bridge Is Finished
Every interview says the center of excellence should eventually step back. None of them say how you know it's time.
That gap matters. "We'll decentralize when the teams are ready" is a feeling dressed as a plan, and feelings don't survive a reorg.
Speero's work on this with Nils Stotz and Paul Drews at Leuphana University, and with Lukas Vermeer of Vista, produced a sharper answer in Aligning Experimentation with Product Operations. Product-driven decentralization has four mechanical preconditions. The experimentation org structures blueprint covers the shapes those preconditions have to hold up. If any one of them is shut, teams cannot own experimentation no matter how the org chart is drawn. A center of excellence that ignores the shut gate calcifies into the gatekeeper everyone in this series warned about.
Architecture. Can a team deploy and test its own changes without waiting on another team? When this gate is shut, velocity is capped by the release train and experiments queue behind unrelated work. This is a hard ceiling. Embed all the experimenters you like into product pods and the central queue simply re-forms somewhere else.
Funding. Is the team permanently funded against outcomes, or project-funded against a fixed scope? Watch for the tell: no budget line exists for a test that loses, and inconvenient results get met with "that's out of scope."
Authority. Can a team change its own spec without approval from outside the team? When this one is shut, test plans get routed to a review board and the center of excellence becomes the bottleneck it was built to remove. A RASCI matrix for experimentation is where most programs discover they never wrote decision rights down.
Culture. Does information flow to where it's needed, or does it get managed for political safety? The symptoms are recognizable anywhere. Results get relitigated, wins get over-claimed, and losses get quietly unshipped.
Reading the gates back against this series explains a pattern that testimony alone leaves unexplained.
RS Group's 95% untested figure and Beerwulf's ignored reports look like discipline problems. They aren't. An experiment is a request to change scope based on evidence, and in a project-funded organization that request is structurally illegal. A team can run a flawless test, read the result correctly, and still have nowhere to put it. That's gate 2, and no amount of rigor opens it. It's also why what a test serves and what it's worth has to be settled before the test runs, not after.
The same logic reframes the gatekeeping failure mode. A center of excellence that reviews and approves test plans is a change approval board for experiments, and there is hard evidence on what those do. The research behind Accelerate found external approval bodies were negatively correlated with lead time, deployment frequency, and time to restore service, while showing no correlation at all with change failure rate.
They slowed delivery down without delivering the safety they exist to provide. Teams using peer review, or no approval process at all, outperformed them.
So Farfetch evolving its peer reviews into structured hypothesis clinics is doing something more specific than sharing best practice. It's replacing sign-off with peer review, which is the arrangement the data actually supports.
One configuration is worth screening for before you celebrate a decentralization. Distributed headcount with centralized approval looks like maturity on an org chart and performs worse than having no structure at all.
Analysts sit inside product pods, the pods still can't act without the center's blessing, and the program gets the coordination cost of decentralization with none of the speed. If you have moved people but not decision rights, you have moved furniture.
The useful thing about framing it this way is that it makes the mandate finishable. "Build a culture of experimentation" never ends. "Open these four gates, then hand over" has a completion state, an owner per gate, and a visible symptom when you're not there yet. Speero's experimentation maturity audit scores the same four areas the gates sit in, benchmarked against 200+ companies.
It's also the cleanest test for whether a transformation is real. Standing up a center of excellence, buying a platform, and reporting test counts every quarter can all happen without a single gate opening. None of the four can be opened by an announcement.
Where This Leaves Your Program
Ten companies, ten different starting points, and the same three-part story. Decentralized testing breaks down. Centralization fixes it and creates new problems of its own.
The fix for those new problems is teaching, standardizing, and getting out of the way.
Layfield, describing how Uber ran things, put a number on what full commitment looks like. "100% of things" got A/B tested there, from backend infrastructure updates to front-end changes. Experimentation was woven into daily operations, not treated as a special project waiting on a JIRA ticket.
That's the state every interview in this series is pointing toward. Some companies call it a Guild. Others call it a Champions Network, or an empowered product team, or nothing at all because it's just how the company works.
If your program is still running everything through a central team, that's likely the right structure for where you are right now.
The work is making sure your experimentation center of excellence doesn't stay that way past the point where it's still solving a real problem.
Read the full series for the details behind every one of these companies. Start with how RS Group rebuilt testing across 31 countries, or go straight to Marty Cagan's case for why optimization alone isn't enough.






















