The Diagnosis: Your QBR scores wins, and that scoring has a failure mode. The ambiguous result gets dressed up as "directionally positive" and lands in the winners column. The clean loss, which told you more, gets one line near the back.
Change the unit of judgment to learning rate, the share of tests at each problem statement that resolved conclusively. A clean loss counts exactly as much as a clean win, and the only outcome that scores zero is the one that returned nothing you can act on.
Rank your problem statements by that number and the quarterly meeting becomes a verdict on whether you picked the right problems.
Your quarterly business review has a scoreboard problem.
It asks what you won last quarter, and that question always has an answer, because there's always a test that moved something and a chart that points up if you pick the window carefully enough.
So the meeting fills with results, everyone nods, and the one question that would change your experimentation strategy next quarter never gets asked. Were these the right problems to work on at all?
One growth team rebuilt the meeting around a different unit of measurement. They started scoring the share of tests at each problem statement that reached a conclusive answer, win or loss.
The goal was a more strategic QBR. They got that. They also got two effects nobody was aiming for, one of which turned up in a different meeting a month later, and one of which changed what the team was able to refuse.
What your quarterly business review actually rewards
Score a quarter on wins and you've built a machine for laundering ambiguous results.
A test comes back with a positive point estimate and a confidence interval straddling zero. Nobody wants to put "inconclusive" on a slide, so it becomes "directionally positive" and goes in the winners column.
Meanwhile a test that came back with a clean, unambiguous loss gets one line near the back, because losses read as failure and the deck is a performance.
You've now rewarded the least informative result in the quarter and buried the most informative one.
McKinsey's read on the QBR is that the meeting usually underdelivers because of what surrounds it, naming unclear ownership, unresolved dependencies, and disconnected OKRs as the culprits. For experimentation programs there's a sixth problem sitting inside the meeting itself, which is that the scoring rewards the wrong thing.
The fix starts by changing what you count.
Learning rate, defined so it can't be gamed
Learning rate is the share of tests at a problem statement that resolved conclusively, measured across the tests that reached their committed runtime.
The word doing the work there is "conclusively," and the whole method collapses if you let that word stay soft. Every test teaches you something if you squint hard enough.
So a test resolves conclusively in exactly 2 situations, both judged at the committed runtime, on the metric you pre-registered before launch.
The primary metric clears zero. The confidence interval around the lift on your pre-registered primary sits entirely above zero. A positive point estimate doesn't qualify, and neither does a p-value that came close. The whole interval clears zero, so the data rules out no-effect and rules out harm. You pre-registered revenue per visitor, the window closes, the interval is +1.2% to +4.8%. That's a win.
A pre-registered cell in your decision matrix fires. Before launch you wrote down how you'd read mixed outcomes. The classic version: if clicks on the CTA drop but revenue per user rises, you call it a win, because you're trading noisy clicks for qualified buyers. If the data lands exactly in that cell, it counts, even though the lead metric went down. It only counts because the cell existed before anyone saw data.
Three things disqualify a result, and they're what make the other two trustworthy.
The wrong metric doesn't count. Your primary was revenue per visitor, add-to-cart went significant, revenue stayed ambiguous. That's not a win. The add-to-cart movement can earn a note against the driver belief, and the test still didn't resolve.
Early peeks don't count. Significant on day 6 of a 21-day commitment is not a result on day 6. Repeatedly checking a running test and stopping when it crosses the line inflates your false positive rate, which Johari and colleagues quantified for online A/B tests at KDD. The stopping rule has to be set before data collection starts, which is the first of the requirements Simmons, Nelson and Simonsohn set out in False-Positive Psychology.
Post-hoc reinterpretation doesn't count. "It worked for mobile users" discovered after the fact is a new hypothesis. Kerr named this pattern HARKing, hypothesizing after the results are known, and described it as presenting a post hoc hypothesis as if it were an a priori one. It's the most common way a program convinces itself it learned something.
All of this is pre-registration, borrowed from clinical trials and imported into a meeting where it has never been enforced.
A clean loss counts as much as a clean win
Losses are the mirror image of wins. The confidence interval sits entirely below zero, or a pre-registered loss cell fires.
And they score identically.
That symmetry is what makes the whole system honest, because it removes the incentive to dress up a shrug. A clean loss tells you exactly as much as a clean win. You now know the driver belief was wrong, you know it at a stated confidence, and you can stop spending on it.
The only result that scores zero is the ambiguous one.
Which inverts what your current QBR rewards. Right now, the ambiguous result gets promoted to a win and the clean loss gets buried. Under learning rate, the ambiguous result is the only actively bad outcome in the quarter, because it consumed traffic and runtime and returned nothing you can act on.
Run that scoring across a quarter and the problem statements sort themselves. The one where 8 of 10 tests resolved is a problem you understand well enough to keep working. The one where 2 of 9 resolved is a problem you've been circling without ever landing a punch.
That second one is the abandon candidate, and until you measured this, it probably looked identical to the first on a results slide.
A note on the name. "Learning rate" measures conclusive outcomes rather than volume of insight, and it collides with the machine learning term of the same name. Speero is standardizing on this definition now, and the previous internal version was looser, so if you adopt it, write the definition down before you use the phrase in front of a client.
There's no universal threshold that separates iterate from abandon. The signal is the ranking across your problem statements rather than an absolute number, which is the same judgment the iterate or move on framework covers at the individual test level.
The division of labour behind a working experimentation strategy
The measurement is half of it. The other half is who does what with it, and this is where the QBR stops being a report.
The growth manager builds the strategic pillars and the problem statements underneath them, drawn from what the program has learned so far and what still matters.
As a reminder, really the main principle behind this is on a quarterly basis being able to determine objectively through the learning rate of each one of the problem statements, are we focusing on the right area? Should we be iterating on that problem statement? Should we abandon it?
Growth manager
That set gets handed to the strategist, who generates hypotheses that ladder up to the problem statements.
The split isn't a hard line, and the team that built it says so. What it does is put the strategic question and the generative question in different hands, at different moments, with a written artifact in between.
Where a program runs customer research alongside testing, the research items come off the same problem statements. One spine, two workstreams.
The pillars themselves work better when you stop guessing them. One growth manager's approach is to ask the client directly what their internal themes are for the coming quarter, rather than inferring them from what's been drip-fed over months. In one case the client handed over 7 themes they were already organized around, and the problem statements wrote themselves from there.
That's a goal tree with a scoreboard attached.
The effect nobody was aiming for
The second effect showed up a month later, in a meeting nobody had connected to the QBR.
The team ran the redesigned QBR in December. In the first week of January, the strategist came back to propose the next batch of test ideas, which is a separate meeting with a separate cadence.
I come back from break first week of January. It really made that for me more focused. I can just come right back to the QBR. I see the strategic pillars and I'm like, oh, I'm just ideating focused in on these. I'm not looking at every swim lane and everything that's going on. I tend to get kind of like cat chasing a laser a little bit with ideation and this really focused it in.
Strategist
Nobody designed the QBR to fix ideation. It fixed ideation anyway, because the output of the quarterly meeting turned out to be the missing input to the monthly one.
The second-order effect goes further. With the strategic framing already settled and written down, the growth manager stops being in the weeds on individual ideas, which is what makes it possible to spread a growth manager across more accounts.
A meeting redesign that changes your staffing model is not a template change.
If your experimentation meeting cadence has a monthly approval rhythm sitting under a quarterly review, this is the handoff that connects them.
What the agreed pillars let you decline
The most useful thing the format does is give you permission to say no, and it works because the client set the terms.
In the December session, the client kept 2 of the 3 proposed themes and replaced the third, because an executive directive had landed that the agency didn't know about. The team took it away, reworked it, and came back a couple of days later.
A month later, that agreement did work nobody had budgeted for.
It kind of just keeps us on track. It's like, remember you guys told us you wanted test ideas here, so we're not going to look at that thing over here. And they're like, oh yeah, right. So it keeps that test approval conversation more focused and aligned.
Strategist
Every experimentation lead has had the meeting where a stakeholder arrives with a pet idea that has nothing to do with the roadmap. An agreed set of pillars turns that from a political conversation into a documented one.
It also cuts the other way.
There is a lot of work in the background, but then you present something that's very succinct, very easy to align on or to say, "No, this is not what we had in mind."
Account lead
The format is designed to make it easier for the client to reject your strategy. That's the feature.
The part that would not automate
Over the next 2 quarters the team automated most of this. The generator pulls ROI, winning tests, the test summary from the quarter's scorecards, and a list of what was learned, and it writes recommendations per problem statement off the learning rate.
Estimates of how far it gets have been given as upwards of 85 to 90%, then around 80 to 90%, then a report of maybe 85 to 90% from one live account. Treat those as one practitioner's estimate on one engagement rather than a measured figure, because no measurement method was ever stated.
The part that stayed manual didn't move across 3 quarters of work.
In March, the forward-looking section was the bulk of the remaining build. In April, the team still refined the strategy section by hand while the tool told them what was working. In July, running live, the framing was that monitoring the output is part of the job.
Three quarters, three rounds of automation, and the same residue every time. The machine can tell you which problem statements resolved and which didn't. Deciding what to do about it next quarter has stayed with a person.
That's worth knowing before you budget for an AI reporting project. The structure of your test reporting automates well. The strategy on top of it doesn't.
What this still doesn't solve
Three open problems, named by the people running it.
The taxonomy isn't locked. There's no agreed language connecting themes to problem statements to individual tests inside the test database, so you can't query your Airtable for every idea related to one problem statement.
We need to lock down that taxonomy, that language there.
Agency leader
That was raised in January and it's still open. If you're building this, solve it at the schema level before you have 200 tests to retrofit.
AI output needs cleanup. The project management function consistently reported that AI-generated strategist output arrives needing tidying before it reaches a client. Automating the draft moves the work rather than removing it.
Nobody has measured whether any of this made the team faster. The honest position from inside the agency is that monthly recurring revenue per strategist is one number with a lot of nuance behind it, and there's no framework yet for evaluating whether the tooling actually moved anything. If you run this in house, your version of that proxy is throughput per person, and it carries the same problem.
What you're adopting is a method that improves focus and traceability, backed by reported reception rather than measured outcomes. Adopt it on that basis.
How to audit your experimentation strategy next quarter
Write down your problem statements before the quarter starts, grouped under 3 to 5 pillars, and get your stakeholders to agree to them in writing.
Pre-register a primary metric and a decision matrix for every test, including the cells where a mixed result still counts. Do it before launch, every time, because the pre-registration is what makes the score mean anything.
At quarter end, score each problem statement on the share of committed-runtime tests that resolved conclusively, counting clean losses equally with clean wins.
Rank the problem statements by that number. Keep the top ones, iterate on the middle, and put the bottom one on the table for abandonment with the evidence in front of you.
Then spend the meeting on the only part that was never going to automate, which is deciding what the next quarter's problems should be.
Your results slide tells you what happened. Your learning rate tells you whether you were pointed at anything worth measuring, and that's the question a strategic testing roadmap exists to answer.
If you want this built into your own program rather than borrowed from ours, that's what our conversion rate optimization work does.























