Experimentation / CRO

Experimentation and Testing Programs acknowledge that the future is uncertain. These programs focus on getting better data to product and marketing teams to make better decisions.

Research & Strategy

We believe that research is an integral part of experimentation. Our research projects aim to identify optimization opportunities by uncovering what really matters to your website users and customers.

Data and Analytics

90% of the analytics setups we’ve seen are critically flawed. Our data analytics audit services give you the confidence to make better decisions with data you can trust.

How AI is Changing Who Builds Experiments and How

Dark blue banner with an orange dot-and-line network graphic and the headline "How AI Is Changing Who Builds Experiments and How," with an "In cooperation with Kameleoon" logo lockup in the top right.

Findings from nineteen long-form interviews with senior experimentation practitioners on the current state and near future of prompt-based A/B testing.

Jonny Longden, Chief Growth Officer, Speero

Research sponsored by Kameleoon

Don't want to read? Watch the 40 minute conversation Jonny recorded with Katie Green at Kameleoon on these findings.

Free Research Report
How AI is Changing Who Builds Experiments: Full Research PDF
Get the complete findings from 19 senior practitioners across 6 industries on prompt-based A/B testing, including all quotes, frameworks, and implementation insights.
LinkedIn is optional but helps us personalize follow-ups
Your PDF is downloading!
If the download doesn't start automatically, click here.
Check your inbox for a confirmation email from Mailchimp to complete your subscription.
19 practitioners
15 min read
PDF
Speero x Kameleoon

01 - Executive summary

The right question is no longer whether AI replaces developers.

AI variation building is the most credible change to hit the A/B testing tool category in a decade. Every practitioner in this research reacted the moment they watched a sentence of plain English turn into a live change on a real page.

The research set out to answer one question: "can AI replace developers in A/B testing?" Nineteen long-form interviews later, that turned out to be the wrong question. The one this paper actually answers is how AI is changing who gets to build a test, and what matters now for the people who used to do that work.

Research at a glance Four stat cards: 19 senior experimentation practitioners interviewed, covering 6 industries (retail, SaaS, travel, fintech, education, healthcare), in 2 hour long-form video sessions hands-on with PBX, including 1 Kameleoon customer with 7 to 8 months of production use. RESEARCH AT A GLANCE INTERVIEWED 19 Senior experimentation practitioners COVERING 6 Industries · retail, SaaS, travel, fintech, education, healthcare FORMAT 2h Long-form video sessions, hands-on with PBX INCLUDING 1 Kameleoon customer with 7 to 8 months of production use speero.com

AI variation tools genuinely expand who can build experiments. They cut the bottleneck that has meant only the best-resourced or most politically connected tests get run. They open up the category of test that would otherwise get cut from a quarterly roadmap.

Engineers shift toward more complex work that genuinely needs them. Variations that win move cleanly into production through capabilities like Kameleoon's PBX Ship.

The bigger question this raises for experimentation as a discipline is about what programs do with the bottleneck that has just been removed, and who gets to participate in the discipline that it used to gatekeep.

02 - What this research was and was not

Nineteen senior practitioners. First-impression evaluations, one power-user case.

Most of the nineteen practitioner sessions were initial-impressions evaluations, not power-user deep dives. Each practitioner spent up to two hours with PBX, built variations on their own sites, prompted it, refined the output, and discussed what they saw.

One of the nineteen, Chris Strobl at Cedars-Sinai, is an existing Kameleoon customer who has been using PBX in production for seven to eight months, including PBX 2.0. His observations are flagged where they come from sustained use rather than first impressions.

Kameleoon released PBX 2.0 during the research window. The practitioners and I worked with the earlier release. This paper flags where 2.0 changes the picture.

Six industries (sectors represented in the research)

Retail · SaaS / Tech · Travel · Fintech · Education · Healthcare

CRO leads · PMs · devs · data scientists

Six industries represented in the research Sectors represented in the research: Retail, SaaS / Tech, Travel, Fintech, Education, Healthcare. Roles interviewed: CRO leads, PMs, devs, data scientists. SIX INDUSTRIES Sectors represented in the research Retail SaaS / Tech Travel Fintech Education Healthcare CRO LEADS · PMs · DEVS · DATA SCIENTISTS FIG. 01 · SECTORS REPRESENTED speero.com

Transparency: this is sponsored research.

Kameleoon paid for the research and provided access to the PBX tool. They asked for an honest read, not a sales piece, and did not intervene in the findings. Where PBX-specific capabilities are described, the paper is explicit. Where category-level observations are made, they apply to all AI variation tools.

03 - How PBX works, in plain terms

A browser extension that reads the live site.

PBX runs as a browser extension. Open a site, type a prompt describing what you want changed, and PBX gathers page context automatically (DOM, CSS, structure, screenshots). It then generates and applies the JavaScript and CSS needed to preview the variant live, and gives you the code.

Fig. 02 — PBX 2.0 rendering a variant on the live site from a plain-English prompt.

The variant is a working version of the actual page, built inside the real context of your site, not a separate mockup built somewhere else. From there, the same variant can be ramped through Kameleoon's platform, complete with feature flagging, audience targeting, and event tracking.

A tool like Claude or Lovable can build a static prototype that looks like your site. But the prototype lives somewhere else and has to be rebuilt to become a test. That is the structural difference that makes the category interesting.

From idea to live test: days or weeks become minutes.

From prompt to live test in one continuous artefact. The variant does not have to be rebuilt to become a test.

The PBX workflow: prompt, context, render, ramp From prompt to live test in one continuous artefact. 01 Prompt: you describe the change, e.g. change the CTA to Find flights. 02 Context: PBX reads the live site — DOM, CSS, structure, screenshots. 03 Render: variant appears on real page — live JS and CSS, code exposed. 04 Ramp: same artefact goes live — feature flagging, audience, tracking. From prompt to live test in one continuous artefact The variant does not have to be rebuilt to become a test. PROMPT 01 You describe the change "Change the CTA to 'Find flights'" CONTEXT 02 PBX reads the live site DOM, CSS, structure, screenshots RENDER 03 Variant appears on real page Live JS + CSS, code exposed RAMP 04 Same artefact goes live Feature flagging, audience, tracking FIG. 03 · THE PBX WORKFLOW speero.com

You describe the change: "Change the CTA to 'Find flights'". PBX reads the live site: DOM, CSS, structure, screenshots. Variant appears on real page: live JS + CSS, code exposed. Same artefact goes live: feature flagging, audience, tracking

Why this matters: the variant is the test, not a mockup of the test.

This is the structural difference between PBX and general-purpose AI coding tools. A prototype built elsewhere has to be rebuilt to become a live experiment. PBX collapses those steps into one continuous artifact.

04 - The benefits practitioners actually care about

Not what the vendor narrative expects.

The benefits practitioners flagged were not raw speed and cost savings. Those came up. The more interesting answers were about access.

The dev queue has got worse, not better.

A consistent theme across every interviewee: getting developer time to build a test is harder now than three years ago, not easier. Every interviewee described the same dynamic, including practitioners at enterprises with significant engineering headcount. Developer bandwidth is constrained everywhere.

"WGU went through a return-to-office mandate. Development resources have become very scarce. To have a test go through the development queue takes quite a long time. Wherever possible I build what I can. If I have to go through dev it can take up to eight weeks." Whitney Norton · Testing and Optimization Manager · Western Governors University

Which tests actually get run.

Every CRO lead I spoke to has a backlog of test ideas that are interesting but not interesting enough to justify a dev sprint. Those tests almost never get run. AI variation tools quietly change that calculation. The cost of trying an idea drops below the threshold at which it needs to be defended.

"I've had many ideas over the years that no one was going to get excited about. Usually I'd need to build up credibility and trust first to convince stakeholders to take on these unique ideas. Now I can do a lot of these tests myself. I don't even have to pitch them half the time." Ben Young · CRO at HubSpot

This is the shape of the shift. Engineers move toward the complex work that genuinely needs them. Practitioners absorb the variation building that should never have required engineering attention in the first place.

Reallocation, not replacement

"In the past, all of the coding was done by the development team. Now, I do most of the coding myself because of AI. The developers only handle the difficult tests." Ben Young · CRO at HubSpot

Stakeholder buy-in is where the immediate value lives.

Practitioners independently described the same use case: showing a stakeholder an idea in a meeting, on their real site, in a state they can interact with. The mockup that wins the room is the same artefact that can be pushed into a live test.

"Having something that you could play with in real time and just say, this is a working project, this is kind of what we're thinking, something actually baked into the site rather than you trying to hack something together. I see that as being really useful." Faith Dallas · Digital Experimentation Manager · New Look

The finding worth carrying: the bottleneck removed is not what most people assume.

Practitioners consistently talked about access, not speed. The mechanical build step getting faster is the surface benefit. The deeper shift is which ideas get to be tests at all, and which people in the organization can participate in experimentation.

05 - What makes this possible: context

The most defensible reason to use PBX over Claude Code.

Multiple practitioners raised, directly or by implication, the question this category will keep facing: why use a specialist tool like PBX when you already have Claude Code, Cursor or GitHub Copilot for everything else? The answer, again and again, was context.

General AI coding tool vs PBX: the context gap Context, the hidden work. General AI coding tools such as Claude Code, Cursor and Copilot require the developer to supply: paste HTML, paste relevant CSS, describe the framework, name the selectors, explain the brand rules — repeated every test. PBX runs against the live site and automatically loads: live DOM as rendered, live CSS in effect, framework and structure, selectors as they exist, brand rules via master prompt in PBX 2.0 — loaded once, reused forever. CONTEXT · THE HIDDEN WORK GENERAL AI CODING TOOL Claude Code · Cursor · Copilot DEVELOPER MUST SUPPLY → Paste HTML → Paste relevant CSS → Describe the framework → Name the selectors → Explain the brand rules Repeated every test. PBX Runs against the live site AUTOMATICALLY LOADED → Live DOM as rendered → Live CSS in effect → Framework and structure → Selectors as they exist → Brand rules via master prompt (2.0) Loaded once. Reused forever. FIG. 04 · GENERAL AI CODING TOOL VS PBX · THE CONTEXT GAP speero.com
"What PBX is trying to do is what we found missing in other tools. The context." Senior Optimisation Manager · Developer background · Major airline

PBX 2.0 extends context into institutional territory.

Kameleoon released PBX 2.0 during the research window. It adds two capabilities that push the tool further into contextual territory: connection to a brand's Figma design modes, and a master prompt where an organization admin encodes design system rules, accessibility constraints, brand voice, and website-specific technical instructions. The tool moves from "faster AI build" toward "an AI build environment that already knows what your site is supposed to look like".

One customer using PBX 2.0 in production described the effect concretely:

"With PBX 2.0, it's got even more capabilities to where it's aware of more of your website, not just the single page. If I have a test I need to run on a dozen pages, it realises, oh, this seems to be kind of like a template. So, do you want me to run this test on any page that is of this template? I'm like, yes, please." Chris Strobl · Cedars-Sinai · Seven to eight months on PBX in production

The ceiling: not the model. What surrounds the model.

The most transferable insight from this research: the limit on what AI variation tools can do sits in the contextual scaffolding around the model, not in the underlying language model itself. Tools that get design system rules, brand voice, accessibility standards, and prior test learnings into the AI's default context will produce visibly better output.

06 - Honest limitations

Where practitioners hesitated.

Trust on the live site.

Practitioners want to inspect and QA any code that could touch live traffic, regardless of who wrote it. PBX exposes the code it produces, supports simulate mode, and allows small-percentage traffic ramps before full release. The build step gets faster. The release process does not change as much.

"If it spits out the code, I would probably put that into a lower environment and test it before putting it into production. I would be curious if the AI code affected something else that a more in-depth QA process would need to catch." Holly Gleason · Digital Testing Manager · US food retailer

Single-page applications.

In the version practitioners tested, SPA frameworks were a limitation (the airline example was Angular). Kameleoon reports that PBX generates code using their JavaScript API (runWhenElementPresent, enableDynamicRefresh) and that the admin master prompt can encode Shadow DOM handling and custom components for the site. A customer running PBX in production on a SPA site described the opposite experience.

"It works quite well, almost surprisingly well. With Target you basically have to bake something into your web pages to inform Target that the page changed. Whereas with Kameleoon, the automatic detection has been great." Chris Strobl · Cedars-Sinai · Sustained production use

Validation for non-developer users.

A data scientist with experience on multiple enterprise experimentation platforms flagged the gap between prompting and debugging. Kameleoon's Configure and Simulate features in PBX 2.0 let users pre-validate test assignment and event firing before launch.

07 - A watch-out for the discipline

Rigour does not come from the build tool.

CRO and experimentation exist as disciplines partly because they impose rigor on what would otherwise be opinion-driven, political or whimsical product decisions. The rigor lives in the hypothesis, the prior research, the success criteria agreed before the test runs, the analysis plan, the statistical thresholds, and the post-experiment write-up, not in the variation itself.

The build tool sits in the middle of this process. It does not produce the rigor around it.

"It's one of those dangerous things that could get out of hand if it's not actually governed." Senior CRO Lead · Large global retailer

The practical translation: reinvest the time. Do not skip.

As AI absorbs variation-building work, make sure the work it gives back is being reinvested in the discipline that surrounds the variation. Better hypotheses, better instrumentation, better analysis. The tools are not designed to protect the rigour for you.

08 - Where this tech is going

The pipeline into production.

Prompt-based building is only the front end of the shift. What comes next is the pipeline that connects a winning variation to the codebase that actually ships.

Kameleoon's PBX Ship extends the workflow into an engineer's IDE via an MCP server. Once a variant wins, the engineer can bring the same artifact into the production repository through PBX Ship, behind a feature flag. The prompt-built variation and the shipped feature are the same object at different stages of the same continuous workflow.

For the wider category, the direction of travel is multi-agent. Different agents handling different parts of the experimentation loop, coordinating through shared context. How cleanly a specialist tool hands off to production is what will separate the tools that keep pace from those that stall as one-off variation builders.

The category signal: prompt to variation to production, in one continuous workflow.

The AI variation tools that endure will be the ones that close the loop between the prompted variant and the shipped feature. That is a solvable problem, and PBX Ship is one credible answer to it already in the market.

09 - So, what is the answer?

Three things AI variation tools actually change.

The right framing is augmentation, not replacement. Across the full range of practitioners, the consistent finding was that these tools change three things.

Three things AI variation tools actually change 01 Who can build a test: the practitioner with an idea is no longer dependent on engineering capacity, design availability and political bandwidth to find out whether the idea is worth pursuing. 02 Which tests get run: second-class ideas that used to die in the backlog now get a chance, the cost of trying one has dropped below the threshold at which it needs to be defended. 03 What specialists do: the work absorbed is the variation-building that should never have required specialist attention, the work amplified is the strategic, judgment-heavy work. 01 Who can build a test. The practitioner with an idea is no longer dependent on engineering capacity, design availability and political bandwidth to find out whether the idea is worth pursuing. 02 Which tests get run. Second-class ideas that used to die in the backlog now get a chance. The cost of trying one has dropped below the threshold at which it needs to be defended. 03 What specialists do. The work absorbed is the variation-building that should never have required specialist attention. The work amplified is the strategic, judgment- heavy work. speero.com

01 · Who can build a test.The practitioner with an idea no longer depends on engineering capacity, design availability, or political capital to find out whether the idea is worth pursuing.

02 · Which tests get run.Second-class ideas that used to die in the backlog now get a chance. The cost of trying one has dropped below the threshold at which it needs to be defended.

03 · What specialists do.The work absorbed is the variation-building that should never have required specialist attention. The work amplified is the strategic, judgment-heavy work.

"For the first time in the sixteen years I've been doing this, I'd be comfortable with more of a marketing type person going in there and doing the prompting." Chris Strobl · Cedars-Sinai · 16 years in experimentation

For experimentation as a discipline, what matters now is what programs do with the bottleneck that has just been removed, and who gets to participate in the discipline it used to gatekeep. Whether AI replaces anyone was never really the question.

What next?

The bigger question is what programs do with the bottleneck that has just been removed.

Try the tool: PBX is free to trial.

Install the browser extension and prompt your first variation on your own site. Start your free trial →

Full article on LinkedIn including practitioner biographies and additional context.

About Kameleoon

Kameleoon is the only prompt-based experimentation platform that enables any team to turn ideas into live A/B tests in minutes, simply by chatting with AI. Built for modern growth teams, Kameleoon adapts to any stack and skill level, offering accurate data, advanced privacy, and native integrations with leading tools. More than 1,000 brands, including Toyota, Mayo Clinic, and Lululemon, trust Kameleoon to scale experimentation without compromising control.

Did you like this article?

(Your feedback helps us write better!) worst 1 - 10 best

Did the article resonate with you?

What aspects did you enjoy or find lacking?
Were there elements you felt we should've covered?
Thank you.
Oops! Something went wrong while submitting the form.

Related Posts

Who's currently reading The Experimental Revolution?