Experimentation / CRO

Experimentation and Testing Programs acknowledge that the future is uncertain. These programs focus on getting better data to product and marketing teams to make better decisions.

Research & Strategy

We believe that research is an integral part of experimentation. Our research projects aim to identify optimization opportunities by uncovering what really matters to your website users and customers.

Data and Analytics

90% of the analytics setups we’ve seen are critically flawed. Our data analytics audit services give you the confidence to make better decisions with data you can trust.

Synthetic research and testing has a place

A green, yellow, and red risk-matrix chart sorting synthetic research by stakes and reversibility, next to the headline 'Synthetic Research and Testing Has a Place' and a photo of author Ben Labay, CEO of Speero.

There is controversy and confusion with synthetic research right now. It's also exploding in applications. Which means this is a fascinating area of work. There is also major legit skepticism, but I mostly think this POV is stuck on the wrong question.

My personal point of view (POV) TL;DR: Synthetic methods can be simultaneously very useful as well as inferior to human-based methods. So....stay calm, both can be true at the same time.

The question everyone keeps fighting over is "is AI-generated data as good as human data?" And the answer is no. Obviously no. It was no last year and it'll be no next year.

But that was never the right question, the same way "is the mouse-study model as good as the human clinical trial" was never the right question in drug development. The right question is: what decision is this data feeding, and what does it cost to be wrong?

That question has different answers in different spots, and the industry is currently terrible at sorting the spots. People like to complain and those complaints get clicks and more attention relative to the other argument that it is truly useful and should get more consideration and experimental applications.

Some context on why I'm writing this. Over the past few months we've been taking demos from companies in this space (Zyent, Listen Labs, Synthetic Users, Squoosh, others). Synthetic panels, simulated shoppers, AI-run interviews, augmented survey samples. The category is exploding, the demos are getting good, and the reaction from the research community has been... hostile.

Some of it is well deserved. Some of it is reflexive. And honestly, this article is partly just an excuse for me to sort out my own position, because I've been holding both reactions at once.

There is the 'content' of the work, and the 'process' of it. Don't confuse the two.

First, the skeptics are right

Let me do the "yes" before the "and," because the criticism of this category is not noise. It's documented, peer-reviewed, and structural.

Speero already made that case in Why I'm Not Sold on Synthetic User Research. That piece still stands. This one is about what the same tools are still good for once you accept they are not users.

MeasuringU reviewed 12 peer-reviewed papers on synthetic users and tallied 14 discouraging findings against 9 encouraging ones. The pattern across them: synthetic outputs often match human averages directionally, then fail when you look at variance. LLMs regress toward the mean of their training data.

They flatten the distribution. They smooth out the outliers, the irrational moves, the edge cases. Which is a problem, because in research the outliers are frequently the finding.

A CHI study by Kapania et al. had 19 qualitative researchers try to recreate their own interview projects with an LLM stand-in. First-pass transcripts looked plausible. Deeper turns fell apart. No contradiction, no palpability, no lived experience underneath the words.

A synthetic user telling you about onboarding frustration isn't remembering anything. It's predicting which words usually sit next to "onboarding frustration."

It gets weirder. The BiasMonkey work tested 9 LLMs for whether they show human survey biases (acquiescence, response order, that family). Mostly they don't. And models tuned hard with RLHF, the alignment process that makes them nice to chat with, acted less human, not more.

They shrugged off the cognitive biases real humans reliably show, then got rattled by typos. Brucks and Toubia found that just reformatting a question or reordering options shifts the results, and that giving a persona a richer backstory paradoxically skews the outputs.

Stack that up and you get the honest summary: an LLM persona is autocomplete wearing a customer's name tag, systematically biased toward the mainstream, English-speaking, tech-literate center of its training data. Using it to represent niche or marginalized populations isn't just inaccurate, it's deeply biased yet disguised not so...which is weird to think about.

And there's a longer-term structural worry called model collapse: as models increasingly train on synthetic content, they progressively lose the tails of the distribution, meaning the rare and weird human traits, which is exactly the stuff research exists to find.

So when researchers push back hard on "synthetic users will replace user research," they're right. Full stop, no push back from me.

And....

The reframe: it's a screen, a filter

None of those failures matter equally everywhere, because not every research task is trying to discover deep human truth. A lot of research work, maybe most of it by volume, is not science rather just filtering.

Which of these 20 ideas deserves real resources? Which survey question is confusing? Which of these 5 layouts is obviously structurally broken?

Drug discovery figured out the right mental model decades ago. You don't run clinical trials on 10,000 compounds. You screen them in silico, simulate them down to a handful, and spend the wet lab and the trial (the expensive, slow, real thing) on the candidates that survived. Nobody in pharma confuses the simulation with the trial.

Nobody skips the simulation either. Biology has the same pattern with model organisms... we learned an enormous amount about human genetics from fruit flies and zebrafish, precisely because everyone stayed clear-eyed that a zebrafish is not a person. The model organism is a cheap, fast proxy that tells you where to point the expensive research.

That's the position synthetic research could (should?) occupy: in front of your real research, never instead of it. A screen, a filter, a pre-flight check. The moment a synthetic result becomes the last word instead of the first pass, you've crossed the line.

And the numbers, for what they're worth right now, support "useful screen" and not "reliable answer." An arXiv pipeline called SimAB ran persona-conditioned agents against 47 historical A/B tests with known outcomes: 67% accuracy overall, 83% on its high-confidence calls. EY used Aaru's agent simulation to replicate a 3,600-person global wealth survey in a day at a 0.90 median correlation with the human data.

Squoosh publishes an 81% match against live test winners on their own benchmark (self-reported...they're grading their own homework there, I haven't pressure tested it, I want to).

On the quant side, a Google Japan pilot with Fairgen improved confidence intervals by 23.8% on average by augmenting small survey cohorts, with caveats: the gains shrink as the base segment grows, it needs clean real seed data to learn from, and it faithfully propagates whatever bias your seed sample carries.

The reframe, by the numbers Four separate benchmark results for synthetic research tools — SimAB, EY with Aaru, Squoosh, and Fairgen — shown as individual stat tiles since accuracy, correlation, match rate, and confidence-interval improvement are not the same unit and should not be charted as one comparable series. The reframe, by the numbers Four studies, four different kinds of metric — not directly comparable SIMAB arXiv pipeline · 47 historical A/B tests 67% accuracy overall 83% accuracy on its high-confidence calls EY / AARU 3,600-person global wealth survey, replicated in a day 0.90 median correlation with human data SQUOOSH own benchmark vs. live test winners 81% match rate CAVEAT Self-reported — their own benchmark, not independently verified. FAIRGEN Google Japan pilot, small survey cohorts 23.8% avg. confidence-interval improvement CAVEAT Gains shrink as the segment grows; needs clean seed data; inherits seed-sample bias. speero.com

Read those numbers as a screen and they're genuinely valuable. 67 to 83% accuracy at pruning your variant list before it touches traffic? Great trade. Read them as an oracle and they're terrifying. A coin you're wrong on 1 time in 3 to 6, deciding your roadmap.

Sorting the category (I see three of em)

Part of why the debate is so muddy is that "synthetic research" is 3 different technologies sold under one label, and I think they fail differently.

  1. Generative emulators. LLM role-play. Give the model a persona, get interview transcripts and reactions (Synthetic Users, Vurvey, Outset live here). Fast, cheap, good at surfacing obvious structural problems. All the qualitative critiques above apply at full strength. It can only remix what's in the training data. It cannot tell you anything genuinely new about your users.
  2. Agentic simulators. Multi-agent populations grounded in real behavioral distributions, census data, transaction data, your analytics (Aaru, Moveo One, Electric Twin, and this is roughly where Squoosh points its synthetic shoppers). These attempt to model what people do rather than what a persona would say, which sidesteps the stated-vs-revealed-preference problem. Output is directional probability, useful for ranking options, never causal proof.
  3. Statistical augmenters. No LLM role-play at all. Math that learns the correlation structure of a real quant dataset and generates additional synthetic rows to firm up thin segments (Fairgen). Narrowest use case, best-validated, and completely dependent on clean seed data.
Sorting the category Three distinct technologies sold under one label — synthetic research — each with its own vendors and its own way of failing: generative emulators, agentic simulators, and statistical augmenters. Sorting the category Three technologies sold under one label — and they fail differently GENERATIVE EMULATORS LLM role-play. Persona in, interview transcripts and reactions out. Fast and cheap at surfacing obvious structural problems. VENDORS Synthetic Users, Vurvey, Outset LIMITATION Can only remix what's already in the training data — it can't tell you anything genuinely new about your users. AGENTIC SIMULATORS Multi-agent populations grounded in real behavioral distributions — census data, transaction data, your own analytics. VENDORS Aaru, Moveo One, Electric Twin (also where Squoosh points its shoppers) LIMITATION Output is directional probability. Useful for ranking options — never causal proof. STATISTICAL AUGMENTERS No LLM role-play. Learns the correlation structure of a real quant dataset, generates rows to firm up thin segments. VENDORS Fairgen LIMITATION Narrowest use case, best-validated of the three — and completely dependent on clean seed data. speero.com

If a vendor can't tell you which of these 3 they are, that's your first red flag.

Where this plugs into Research

Now the part where I connect this to how we actually work, because abstract frameworks are cheap.

Our research practice at Speero runs on ResearchXL (the model Peep Laja developed back in the CXL days): heuristic analysis, technical analysis, digital analytics, qualitative research (surveys, interviews), user testing, session replays.

Six evidence streams, triangulated. The whole point of the model is that no single data source gets trusted alone. Every finding needs corroboration from another stream before it drives a test.

Here's what I've realized sorting through these demos: synthetic research is not a seventh stream. It doesn't earn a pillar, because it doesn't produce ground truth about your users. What it does is make the existing six streams cheaper to run well. It's a throttle on the front of each pillar, not a new pillar:

  • Before qualitative: pilot your interview script and survey instrument on synthetic respondents. If the model misreads your question, humans will too. You just saved a fielding round.
  • Before user testing: run synthetic agents through the new checkout flow to catch the glaring friction (broken logic, absurd cognitive load) so your expensive human sessions are spent on nuance, not on discovering the obvious.
  • Before A/B testing: prune. The build cost of variants collapsed with AI, so backlogs are exploding while traffic stayed exactly the same. Something has to cut the list, and today that something is usually seniority in the room. A simulated pre-read is a better cutter than a HiPPO, even at 67% accuracy.

And the triangulation/synthesis habit ResearchXL drilled into us turns out to be exactly the immune system synthetic data requires. A synthetic finding is a hypothesis. It gets corroborated by a real stream or it dies. Teams that already operate this way can absorb synthetic tools safely almost by default.

Teams that were already prone to trusting single data sources are going to hurt themselves fast, because synthetic data is the most convincing single source ever built. It always answers. It's never confused. It never says "weird, let me think about that." Real humans (and real human responses) are messy.

The growth experimentation lens says the same thing from a different angle. If you run your program like a portfolio of bets (I've written about this before), then live traffic is your investable capital. It's finite, it's the most expensive resource in the whole lifecycle.

Synthetic research is cheap diligence on which bets deserve funding. No fund skips diligence because diligence is sometimes wrong.

The three zones

Getting geeky and taxonomizing for a second, because this is the actual positioning framework. Sort any proposed use of synthetic data by two variables: the stakes of the decision, and the reversibility of being wrong.

The three zones A diagonal risk matrix sorting synthetic research use cases by the stakes of the decision and the reversibility of being wrong: a green zone for cheap, reversible uses, a yellow zone for directional uses that must be triangulated, and a red zone for uses synthetic data should never drive. The three zones Sorted by the stakes of the decision and the reversibility of being wrong RED YELLOW GREEN REVERSIBILITY OF BEING WRONG → low high STAKES OF THE DECISION → low high GREEN ZONE cheap to be wrong, fast to reverse Pre-flight friction checks, variant pruning, piloting research instruments, and statistical augmentation with a real holdout. YELLOW ZONE directional, must be triangulated Scenario planning, pricing sensitivity ranking, and proto-personas for markets with zero data. RED ZONE do not Final validation, anything requiring lived experience, and research on vulnerable or marginalized populations. speero.com

Green zone (cheap to be wrong, fast to reverse). Pre-flight friction checks on flows and forms. Variant pruning before live tests. Piloting research instruments before fielding them. Statistical augmentation of clean, closed-ended quant data with a real holdout for validation. In this zone the cost of the AI being wrong is a few API credits and a slightly worse shortlist. Use it aggressively.

Yellow zone (directional, must be triangulated). Market scenario planning, pricing sensitivity ranking, proto-personas when you're entering a market where you have zero data and need a temporary scaffold. Everything here is a falsifiable hypothesis, labeled as such, with an explicit plan for real data to overwrite it. If a yellow-zone output survives more than one quarter without human corroboration, that's a process failure.

Red zone (do not). Final validation of a feature, a campaign, or a pricing change. Anything requiring lived experience, emotional truth, or cultural nuance. Any research on niche, vulnerable, or marginalized populations. And replacing foundational qualitative research entirely, which just builds an echo chamber that confirms whatever your org already believed. In the red zone the only acceptable data is real behavior: a live experiment or actual purchases.

Two operating rules ride along with the zones. First, label everything. Synthetic data gets marked as synthetic in every report, deck, and dashboard, no exceptions (ESOMAR's guidance is blunt on this, and the updated ICC/ESOMAR Code now covers synthetic data explicitly, and they're right: presenting model output as human response is where the credibility of the whole research function dies...should think the same with AI output...like half this post!).

Second, mind your calibration data. These tools are only as good as the behavioral data grounding them, and if your analytics run on client-side tracking, ad blockers and privacy browsers are already eating a meaningful chunk of your real sessions (often a quarter or more, and worse for tech-savvy audiences) before bot traffic even enters the picture. Garbage in, confident-sounding garbage out.

So what about the outliers?

The learnings that matter most in experimentation are the ones where people don't behave the way you'd expect. Flat results. Backwards results. A model trained on what usually happens will be excellent at predicting what usually happens.

How does it ever catch the weird ones? Because if it only confirms priors, it's not a research tool, it's a comfort blanket.

My honest guess is that today's tools catch some structural surprises but miss most behavioral ones, and that the vendors' own benchmarks aren't built to expose the difference. Someone should build that benchmark. Maybe we will. Likely.

Where I land, for now: the absolutists on both ends are making the same mistake, which is treating "synthetic research" as one thing that's either valid or invalid. ATM it's a screen. Screens are judged by what they let through to the expensive stage, not by whether they replace it. Put it in front of your real research, label it, triangulate it.

Are you using any of this stuff yet? I'm especially curious about anyone who's run a synthetic pre-read against a live test and watched it be wrong. Those stories are worth more than the benchmarks right now.

Did you like this article?

(Your feedback helps us write better!) worst 1 - 10 best

Did the article resonate with you?

What aspects did you enjoy or find lacking?
Were there elements you felt we should've covered?
Thank you.
Oops! Something went wrong while submitting the form.

Related Posts

Who's currently reading The Experimental Revolution?