Experimentation / CRO

Experimentation and Testing Programs acknowledge that the future is uncertain. These programs focus on getting better data to product and marketing teams to make better decisions.

Research & Strategy

We believe that research is an integral part of experimentation. Our research projects aim to identify optimization opportunities by uncovering what really matters to your website users and customers.

Data and Analytics

90% of the analytics setups we’ve seen are critically flawed. Our data analytics audit services give you the confidence to make better decisions with data you can trust.

How to Talk to Clients About AI-Generated Test Ideas

The Diagnosis: I built a set of test concepts with an agent that checks them against a client's strategic pillars, problem statements, and every test they've run before. Twenty minutes instead of two days. Then I presented the concepts as my own recommendations, and the room went quiet, the specific quiet of someone who doesn't want to say no to a person they like.
Rejecting a personal recommendation costs a client something social, so they hesitate to say no even when a concept is wrong for the account. Naming the agent and framing the set as a first pass moves that cost to zero.
Swap one sentence in how you introduce those concepts, and the same data gets a different reception in the room.

I built a set of test concepts with an agent that checks them against the client's strategic pillars, their open problem statements, and every test they've run before. Twenty minutes instead of two days.

Then I got on the call and presented the concepts as my own recommendations. The room went quiet, the specific quiet of someone who doesn't want to say no to a person they like.

Then I noticed something changes when I swap one sentence in that pitch. Same concepts, same data.

The only thing that moves is how I introduce where they came from. Here's the sentence, why it works, and where the evidence for it stops.

What "Vetted" Actually Means

That sentence leans on one word: vetted. If I call a raw brainstorm dump "your top five opportunities," that framing holds right up until a client catches a concept that contradicts something the account already tried.

Then it collapses and takes trust with it. So before the sentence itself, it's worth being specific about what MILES, our CRO ideation agent, does to earn that word.

MILES diagnoses the page before it generates a single idea, and we hold that diagnosis to a specific bar. A note that only restates the obvious doesn't clear it.

Our documentation for MILES spells out the difference with a worked example. "The CTA could be clearer" is a thin diagnosis, the kind MILES is built to reject.

Here's what a sharp one looks like instead: a primary CTA sitting below several secondary CTAs of near-identical visual weight, with usability research showing visitors scanning that cluster for a few seconds before a meaningful share click the wrong button.

A designer can act on that version without a follow-up question. Nobody can act on "the CTA could be clearer."

Getting there means checking the page from a few angles before a single idea gets generated: what the visitor's actually being asked to decide, whether the page's hierarchy matches what they need first, and whether the trust signals are specific enough to survive scrutiny instead of filler like "trusted by thousands." Friction matters too, where it sits between interest and action, but it's usually the smallest of the four.

Every research finding ties to a named page element and two or three possible mechanisms. Not filed away as a general theme.

Only once a diagnosis clears that bar does MILES start generating ideas. Even then, it produces 15 to 25 raw concepts across different angles before anyone starts picking favorites.

AI-Generated Test Ideas Aren't the Trust Problem

Client pushback on visibly AI-assisted work isn't the trend you'd expect. I tracked it across my own accounts in Q1 2026: clients are using more AI on their own end, and their resistance to AI-generated content is trending down, not up.

A few of my accounts have even asked us to run workshops on how we use AI internally.

The resistance I run into on client calls is approval psychology, not distrust of the tool. When I present a set of test concepts as "these are the exact concepts we're going to run, and I came up with them," the client hears something specific.

A "no" means rejecting a person, not an idea. Clients maybe feel a little more able to say no when the concepts don't read as my personal work, and worry less about hurting my feelings.

That's a harder problem than skepticism about AI. What a "no" costs the client is social.

It has nothing to do with whether they trust the technology behind the recommendation.

The Reframe: Presenting AI-Generated Test Ideas as a First Pass

Here's the actual script I delivered at our April all-hands:

"This went through MILES. MILES is this agent that cross references our current strategic pillars, the problem statements, all of your past test data, and then vetted it as these are your top five opportunities to consider, and I did a first pass, but let's go through these together."

Three things are doing work in that sentence. It names the agent instead of hiding the process.

It states exactly what I checked against: the pillars, the problem statements, the test history, so the client knows the recommendation isn't a guess. And it ends by handing the client a seat at the table.

"Let's go through these together" is a very different sentence from "here's what we're running."

I pair the script with a specific delivery mechanic. I pull the concepts up in a tool like Tessa, where I can show the actual screenshots and reorder or cut ideas live as I talk.

That turns the review from a defense of finished work into a working session. The client helps shape a set of concepts that's already most of the way there, instead of sitting in judgment of a finished deliverable.

I want to flag the limits of my own claim here. I'd call the softer-rejection effect purely anecdotal, not across all clients, based on what I've personally noticed rather than anything measured.

Treat it as a hypothesis worth testing on your own accounts, not a settled result.

The Reframe: Old Script vs. New Script A comparison of two ways to present AI-generated test concepts to a client. The old script presents finished concepts as the strategist's personal work, which makes a "no" feel like rejecting a person. The new script names the agent MILES, states what it checked, and frames the concepts as a first pass to review together, which makes a "no" cost the client nothing socially. OLD SCRIPT "These are the exact concepts we're going to run." "I came up with them." No visibility into what was checked before the call. RESULT A "no" feels like rejecting you personally. NEW SCRIPT "This went through MILES." Checked against your pillars, problem statements, and test history. "I did a first pass. Let's go through these together." RESULT A "no" costs nothing. You're shaping the work, not rejecting a person. speero.com

How Weak Ideas Get Cut

Diagnosis produces 15 to 25 raw concepts, not a final list. Turning that into something worth calling vetted is where the filtering happens.

Every surviving idea gets checked against the client's exact strategic pillars and problem statements, not a paraphrase of them. Weak ideas get cut for one of three reasons.

The mechanism is too indirect to plausibly change behavior. The idea targets a footnote instead of a real bottleneck.

Or it's too vague for a designer to build without asking follow-up questions.

What survives gets graded on evidence strength: Strong, Moderate, or Weak. A Strong idea has multiple sources (test history, research, and diagnostics) all pointing the same direction.

A Weak idea rests on a single data point, or a general behavioral principle with no client-specific backing behind it. Our rule is that at least 60% of a final set has to rate Moderate or Strong.

If an account's data can't support that ratio, the agent hands back a shorter list instead of padding it with weak ideas to hit a round number of five.

The MILES Pipeline: From Client Data to Vetted Opportunities Six stages, arranged in two rows of three, showing how MILES turns client data into vetted test concepts: strategic context, diagnose, ideate, filter and grade, duplication check, and five to seven vetted concepts. THE MILES PIPELINE: FROM CLIENT DATA TO VETTED OPPORTUNITIES 1. STRATEGIC CONTEXT Client pillars, problem statements, and test history from live account systems. 2. DIAGNOSE Names the page element, user behavior, and evidence before ideation starts. 3. IDEATE Generates 15 to 25 raw concepts across different angles, unfiltered. 4. FILTER & GRADE Cuts weak mechanisms. Grades survivors Strong, Moderate, or Weak. 5. DUPLICATION CHECK Checks against tests already running or concluded on the same page. 6. 5-7 VETTED CONCEPTS What reaches the client as your top opportunities to consider. speero.com

The last check happens before anything reaches a deck. MILES runs a duplication pass against the account's live test history, catching concepts that collide with a test that's already running or already concluded on the same page.

That step exists because a recommendation that repeats something the account tried six months ago is exactly what makes "vetted" sound like a sales word instead of a description.

None of that removes the strategist from the loop. Around the same time I noticed the softer-pushback trend, our project managers flagged that AI output from strategists still needs cleanup before it reaches a client.

The script above is a framing choice, not a shortcut around review. A strategist still has to read what MILES produced, check it against the account, and be ready to defend any concept a client pushes on.

That defense should point to the actual pillar, problem statement, and data strength rating behind the idea, not an improvised justification.

The Agent Isn't Supposed to Be Finished

Part of what makes the script credible is being honest that MILES is a work in progress. By April, I estimated the agent had been through 20 to 30 versions, each one a response to something in its output that wasn't working.

My point to the wider strategy team was about the expected shape of building with these tools, not something to apologize for.

"It's okay to just get something out there and have the output not be super great. We need to start somewhere."

That's worth saying to clients too, in the right moment, because it reinforces the same framing as the script itself. An agent that's openly still being tuned is easier to treat as a collaborator you're both evaluating than an oracle whose verdict you either accept or reject.

The first-pass framing and the still-in-development framing are the same move, made twice.

What Changes in the Room

The deck stays the same either way. The pillars, the problem statements, the five concepts, and the supporting data are identical whether you present them as your personal picks or as MILES's first pass.

What shifts is what the client is being asked to do when they push back.

Rejecting your first pass on a framework you've both just walked through together costs the client nothing socially. Rejecting your personal recommendation does.

If you run experimentation programs and you've been presenting AI-assisted concepts as your own finished work, there's a cheap way to test this.

Change one sentence on your next call and watch what happens to the conversation. It costs nothing extra to say, and it's the only way to find out if it moves your client relationships the way it moved mine.

Did you like this article?

(Your feedback helps us write better!) worst 1 - 10 best

Did the article resonate with you?

What aspects did you enjoy or find lacking?
Were there elements you felt we should've covered?
Thank you.
Oops! Something went wrong while submitting the form.

Related Posts

No items found.

Who's currently reading The Experimental Revolution?