The Diagnosis: Mature experimentation programs don't stall from a lack of ideas. They stall because years of test history and research sit in decks nobody reopens.AI can turn that backlog into first-draft hypotheses in minutes, but only if you build the memory layer first. Skip that step and you're automating guesswork faster, not accelerating ideation.
In experimentation, the longer a program runs, the more you realize that test volume alone isn't the goal. Learning velocity is, too. But for mature programs that have run dozens or even hundreds of tests, keeping track of what you've already learned, and how to build on it, is increasingly difficult.
Across many programs I've worked on, a consistent pattern emerges: over time, it becomes harder to synthesize years of insights. Strategists and PMs alike find themselves wondering, "Did we already test this?" or "What were the findings from that survey?" And when it's time to ideate, we risk rehashing old ideas instead of building on past knowledge.
In my experience, this is a strategic challenge, not a tooling one. And it's where AI, thoughtfully applied, becomes a force multiplier for strategic memory and momentum, not a shortcut.
The Strategic Bottleneck in Mature Programs
At earlier stages of an experimentation program, the ideation process often feels fresh. New ideas emerge quickly, and it's easy to identify gaps. But once a program hits maturity, spanning multiple teams, pages, and years, the sheer volume of knowledge becomes a barrier in itself.
Programs rarely stop testing outright. The real risk is testing in isolation. That risk shows up in the data too.
VWO's Experimentation Maturity Benchmark 2025-26 found that 56% of experimentation programs are stuck in an execution bottleneck, and for mature programs that bottleneck is rarely about running more tests. It's about knowing what's already been learned. Valuable insights from past wins (and losses) get buried in static decks or scattered docs.
And even with well-maintained documentation systems, surfacing relevant learnings during ideation can feel like an uphill climb. This is where AI becomes necessary. When used well, it becomes a bridge between memory and creativity.
Build a System of Record for Learning
Tools: Airtable + Airtable Omni, NotebookLM
The foundation of a strong ideation loop starts with having a centralized, queryable system of record. For the testing programs we run at Speero, Airtable is the ideal base, serving as the strategic and operational source of truth for experiment tracking, metadata, and insights.
And if Airtable Omni is enabled, even better, you can begin querying that test memory directly. For example:
- "Show me every test we've run on the PDP targeting image zoom behavior"
- "What insights have we gathered about mobile nav users over the past 6 months?"
These queries can be incredibly powerful when Omni is layered on top of well-structured data. That said, for more unstructured or qualitative insight, such as heatmap observations, moderated interview takeaways, or workshop slides, I've found it extremely helpful to create a parallel repository using NotebookLM.
In my workflow, I usually maintain two distinct notebooks:
- Testing Repository: A backup of all test outcomes (especially helpful if Airtable isn't yet enabled with Omni or if multiple formats are used)
- Research Repository: Loaded with interview notes, survey results, messaging testing, usability studies, etc.
NotebookLM's strength lies in pulling thematic insights from more narrative data sources, especially when those insights are tucked away in slide decks, PDFs, or multi-page documents. While Airtable Omni excels at structured recall, NotebookLM helps unearth more nuanced behavioral signals.
If you have both set up, you're in a very strong position: Airtable for structured test memory, NotebookLM for qualitative deep dives. The two complement each other.
Seed the Strategic Prompt
Tool: Gemini (custom Gem), or any LLM
Once I've pulled together relevant insights, either from Omni or NotebookLM, I shift into idea generation mode using a custom Gemini Gem designed specifically for test ideation.
You can use any LLM here (Claude, ChatGPT, etc.), but the Gemini Gem I use is tailored to accept:
- Thematic insights from past research and testing
- A user context (e.g. PDP shopper, logged-in dashboard user)
- A screenshot or brief of the page or feature being targeted
- Known friction areas or business KPIs to prioritize
Here's a typical prompt flow:
"Given what we know about PDP users and their color selection behaviors below, suggest 4-5 test ideas that could improve confidence in selecting the right variant, particularly on mobile."
You're creating a starting point for creative iteration, not a perfect idea on the first try.
I'll often follow up with:
- "We've already tested tooltip-based solutions, what else might work?"
- "Can you suggest ideas that require minimal dev effort?"
- "What's a high-risk/high-reward concept for this user segment?"
This is where the AI starts becoming a collaborator, surfacing ideas you may not have considered, reframing the problem, or pushing you into lateral directions you can build from.
Human Synthesis and Creative Close
The key to making this process valuable is knowing where AI ends and strategy begins.
AI can't see around corners. It doesn't understand internal roadmaps, political friction, or the nuance behind a failed test from 9 months ago. In my experience, what it can do is get you well past the first draft, fast. Then you bring it home.
This means:
- Stress-testing the hypothesis: Is it clear, testable, and aligned with a known problem statement?
- Scoping for feasibility: Would dev even build this?
- Connecting back to themes: Does it help us validate or refute a bigger behavioral pattern?
One of the biggest mental shifts I've made is realizing that ideation is about relevance. Most people chase originality instead. The best test ideas are rarely surprising. They're timely, contextual, and built on what we've already learned.
Visualize to Align
Tool: Replit, Figma
Once you've got a strong test concept, bringing it to life visually can help build alignment across stakeholders and accelerate buy-in. While we'll often use Figma for high-fidelity prototypes, I've found Replit to be incredibly useful for quick, low-lift mockups.
If we're pitching a new mobile PDP variation, I'll often create a quick mock showing the proposed change, maybe it's a color selector with embedded visual guidance or a new sticky behavior on scroll. A rough version that conveys the intent is enough to get the idea out of a doc or slide deck.
From Test Velocity to Strategic Continuity
When programs reach scale, the bottleneck shifts from test velocity to strategic continuity. AI helps solve that by amplifying our access to what we already know and helping us move from insight to execution faster.
This is the future of ideation in experimentation:
- A memory layer powered by Airtable + Omni
- A synthesis layer powered by NotebookLM
- An ideation layer powered by Gemini
- A visualization layer powered by tools like Replit
And at the center of it all, the strategist. Still asking the right questions, refining the noise, and pushing for clarity.






















