Experimentation / CRO

Experimentation and Testing Programs acknowledge that the future is uncertain. These programs focus on getting better data to product and marketing teams to make better decisions.

Research & Strategy

We believe that research is an integral part of experimentation. Our research projects aim to identify optimization opportunities by uncovering what really matters to your website users and customers.

Data and Analytics

90% of the analytics setups we’ve seen are critically flawed. Our data analytics audit services give you the confidence to make better decisions with data you can trust.

How To Use AI for Better Ideation in Mature Experimentation Programs

Blog post header: 'How to Use AI for Better Ideation in Mature Experimentation Programs,' with Speero logo, headshot of Alexander Loesch, Director of Strategy at Speero, and a preview of four framework layers including memory, synthesis, ideation, and visualization tools.
The Diagnosis: Mature experimentation programs don't stall from a lack of ideas. They stall because years of test history and research sit in decks nobody reopens.AI can turn that backlog into first-draft hypotheses in minutes, but only if you build the memory layer first. Skip that step and you're automating guesswork faster, not accelerating ideation.

In experimentation, the longer a program runs, the more you realize that test volume alone isn't the goal. Learning velocity is, too. But for mature programs that have run dozens or even hundreds of tests, keeping track of what you've already learned, and how to build on it, is increasingly difficult.

Across many programs I've worked on, a consistent pattern emerges: over time, it becomes harder to synthesize years of insights. Strategists and PMs alike find themselves wondering, "Did we already test this?" or "What were the findings from that survey?" And when it's time to ideate, we risk rehashing old ideas instead of building on past knowledge.

In my experience, this is a strategic challenge, not a tooling one. And it's where AI, thoughtfully applied, becomes a force multiplier for strategic memory and momentum, not a shortcut.

The Strategic Bottleneck in Mature Programs

From memory to momentum: the four-layer AI ideation loop Four layers: 01 Memory layer, Airtable plus Omni, structured test memory queryable on demand. 02 Synthesis layer, NotebookLM, thematic insight pulled from qualitative narrative sources. 03 Ideation layer, Gemini or any LLM, fast first-draft test concepts from context and prompts. 04 Visualization layer, Replit and Figma, quick mockups that build stakeholder alignment. At the center of it all: the strategist, still asking the right questions, refining the noise, pushing for clarity. THE AI-ASSISTED IDEATION LOOP From Memory to Momentum Four layers carry the work. One strategist keeps it honest. 01 MEMORY LAYER Airtable + Omni Structured test memory, queryable on demand. 02 SYNTHESIS LAYER NotebookLM Thematic insight pulled from qualitative, narrative sources. 03 IDEATION LAYER Gemini (or any LLM) Fast first-draft test concepts from context and prompts. 04 VISUALIZATION LAYER Replit + Figma Quick mockups that build stakeholder alignment. AT THE CENTER OF IT ALL: THE STRATEGIST Still asking the right questions. Refining the noise. Pushing for clarity. FIG. 01 · THE FOUR-LAYER IDEATION LOOP speero.com

At earlier stages of an experimentation program, the ideation process often feels fresh. New ideas emerge quickly, and it's easy to identify gaps. But once a program hits maturity, spanning multiple teams, pages, and years, the sheer volume of knowledge becomes a barrier in itself.

Programs rarely stop testing outright. The real risk is testing in isolation. That risk shows up in the data too.

VWO's Experimentation Maturity Benchmark 2025-26 found that 56% of experimentation programs are stuck in an execution bottleneck, and for mature programs that bottleneck is rarely about running more tests. It's about knowing what's already been learned. Valuable insights from past wins (and losses) get buried in static decks or scattered docs.

And even with well-maintained documentation systems, surfacing relevant learnings during ideation can feel like an uphill climb. This is where AI becomes necessary. When used well, it becomes a bridge between memory and creativity.

Build a System of Record for Learning

Two memory layers, two jobs Airtable plus Omni: structured, queryable test memory covering experiment tracking, metadata and insights, queried directly via Omni, for example show me every test on the PDP. NotebookLM: qualitative synthesis via a Testing Repository backing up outcomes and a Research Repository of interviews, surveys and workshops, surfacing themes buried in decks and PDFs. The two complement each other. BUILD A SYSTEM OF RECORD Two Memory Layers, Two Jobs Structured recall from Airtable. Thematic synthesis from NotebookLM. AIRTABLE + OMNI Structured, queryable test memory WHAT IT DOES → Experiment tracking, metadata, insights → Query structured data directly via Omni → "Show me every test on the PDP" Best for structured recall. NOTEBOOKLM Qualitative synthesis across research WHAT IT DOES → Testing Repository — backup of outcomes → Research Repository — interviews, surveys, workshops → Surfaces themes buried in decks & PDFs Best for thematic deep dives. "The two complement each other." FIG. 02 · STRUCTURED RECALL VS QUALITATIVE SYNTHESIS speero.com

Tools: Airtable + Airtable Omni, NotebookLM

The foundation of a strong ideation loop starts with having a centralized, queryable system of record. For the testing programs we run at Speero, Airtable is the ideal base, serving as the strategic and operational source of truth for experiment tracking, metadata, and insights.

And if Airtable Omni is enabled, even better, you can begin querying that test memory directly. For example:

  • "Show me every test we've run on the PDP targeting image zoom behavior"
  • "What insights have we gathered about mobile nav users over the past 6 months?"

These queries can be incredibly powerful when Omni is layered on top of well-structured data. That said, for more unstructured or qualitative insight, such as heatmap observations, moderated interview takeaways, or workshop slides, I've found it extremely helpful to create a parallel repository using NotebookLM.

In my workflow, I usually maintain two distinct notebooks:

  • Testing Repository: A backup of all test outcomes (especially helpful if Airtable isn't yet enabled with Omni or if multiple formats are used)
  • Research Repository: Loaded with interview notes, survey results, messaging testing, usability studies, etc.

NotebookLM's strength lies in pulling thematic insights from more narrative data sources, especially when those insights are tucked away in slide decks, PDFs, or multi-page documents. While Airtable Omni excels at structured recall, NotebookLM helps unearth more nuanced behavioral signals.

If you have both set up, you're in a very strong position: Airtable for structured test memory, NotebookLM for qualitative deep dives. The two complement each other.

Seed the Strategic Prompt

Tool: Gemini (custom Gem), or any LLM

Once I've pulled together relevant insights, either from Omni or NotebookLM, I shift into idea generation mode using a custom Gemini Gem designed specifically for test ideation.

You can use any LLM here (Claude, ChatGPT, etc.), but the Gemini Gem I use is tailored to accept:

  • Thematic insights from past research and testing
  • A user context (e.g. PDP shopper, logged-in dashboard user)
  • A screenshot or brief of the page or feature being targeted
  • Known friction areas or business KPIs to prioritize

Access the Gemini Gem here

Here's a typical prompt flow:

"Given what we know about PDP users and their color selection behaviors below, suggest 4-5 test ideas that could improve confidence in selecting the right variant, particularly on mobile."

You're creating a starting point for creative iteration, not a perfect idea on the first try.

I'll often follow up with:

  • "We've already tested tooltip-based solutions, what else might work?"
  • "Can you suggest ideas that require minimal dev effort?"
  • "What's a high-risk/high-reward concept for this user segment?"

This is where the AI starts becoming a collaborator, surfacing ideas you may not have considered, reframing the problem, or pushing you into lateral directions you can build from.

Human Synthesis and Creative Close

Where AI stops and strategy begins Three checkpoints for human synthesis after AI ideation: 01 Stress-test the hypothesis — is it clear, testable, and aligned with a known problem statement? 02 Scope for feasibility — would dev even build this? 03 Connect back to themes — does it help us validate or refute a bigger behavioral pattern? HUMAN SYNTHESIS Where AI Stops and Strategy Begins AI gets you 60% of the way there. This is how you bring it home. 01 Stress-test the hypothesis Is it clear, testable, and aligned with a known problem statement? 02 Scope for feasibility Would dev even build this? 03 Connect back to themes Does it help us validate or refute a bigger behavioral pattern? FIG. 03 · THE HUMAN SYNTHESIS CHECKPOINT speero.com

The key to making this process valuable is knowing where AI ends and strategy begins.

AI can't see around corners. It doesn't understand internal roadmaps, political friction, or the nuance behind a failed test from 9 months ago. In my experience, what it can do is get you well past the first draft, fast. Then you bring it home.

This means:

  • Stress-testing the hypothesis: Is it clear, testable, and aligned with a known problem statement?
  • Scoping for feasibility: Would dev even build this?
  • Connecting back to themes: Does it help us validate or refute a bigger behavioral pattern?

One of the biggest mental shifts I've made is realizing that ideation is about relevance. Most people chase originality instead. The best test ideas are rarely surprising. They're timely, contextual, and built on what we've already learned.

Visualize to Align

Tool: Replit, Figma

Once you've got a strong test concept, bringing it to life visually can help build alignment across stakeholders and accelerate buy-in. While we'll often use Figma for high-fidelity prototypes, I've found Replit to be incredibly useful for quick, low-lift mockups.

If we're pitching a new mobile PDP variation, I'll often create a quick mock showing the proposed change, maybe it's a color selector with embedded visual guidance or a new sticky behavior on scroll. A rough version that conveys the intent is enough to get the idea out of a doc or slide deck.

From Test Velocity to Strategic Continuity

When programs reach scale, the bottleneck shifts from test velocity to strategic continuity. AI helps solve that by amplifying our access to what we already know and helping us move from insight to execution faster.

This is the future of ideation in experimentation:

And at the center of it all, the strategist. Still asking the right questions, refining the noise, and pushing for clarity.

Did you like this article?

(Your feedback helps us write better!) worst 1 - 10 best

Did the article resonate with you?

What aspects did you enjoy or find lacking?
Were there elements you felt we should've covered?
Thank you.
Oops! Something went wrong while submitting the form.

Related Posts

Who's currently reading The Experimental Revolution?