Experimentation / CRO

Experimentation and Testing Programs acknowledge that the future is uncertain. These programs focus on getting better data to product and marketing teams to make better decisions.

Research & Strategy

We believe that research is an integral part of experimentation. Our research projects aim to identify optimization opportunities by uncovering what really matters to your website users and customers.

Data and Analytics

90% of the analytics setups we’ve seen are critically flawed. Our data analytics audit services give you the confidence to make better decisions with data you can trust.

From Agent Chain to Harness: Speero's AI Skills Library

Article banner for "From Agent Chain to Harness: Speero's AI Skills Library" by Alexander Loesch, Director of Strategy at Speero, showing a headshot of Loesch beside cards for the agent chain's roles, Dexter, Miles, and Brixton on ideation and strategy, Blast and Rema on orchestration.
The Diagnosis: A chain of named AI agents can work, and still be the wrong thing to keep building. On May 19, Speero's Dexter-Miles-Brixton chain caught three of six new test concepts as duplicates before anyone had asked it to check.
Three months later, the leadership that built that chain called it "advanced context engineering," a level to climb past rather than a finished system. The chain that sold the room in May was already the old way of working by August.
Measure your own skills library against that same test, and the real question becomes whether you're managing a chain of skills, or the harness they run inside.

If you're building a skills library of your own right now, the version you ship this quarter will embarrass you by next quarter. That's the actual shape of the work, and our own internal AI system is the proof.

In May, I handed an AI agent nothing but an Airtable base ID and watched it figure out the rest on its own. It read the schema without being told the field names. It matched six new test ideas against what was already sitting in the client's backlog. Then it stopped and said three of them already existed.

Nobody had asked it to check for duplicates. It checked anyway.

That moment is the reason this piece exists. It's also why it almost didn't get written as one piece. By August, the same leadership I'd built that chain with stood in front of the whole agency and said it was already the old way of working. Here's what happened in between, and what it means for your own skills library.

The chain that sold the room

On May 19, I ran a live demo for the agency's weekly all-hands, and opened with the plain version of what I'd built:

...The TLDDR is that we have three agents. Dexter, Miles, and Brixton. Dexter synthesizes the strategic pillars and problem statements for a program by looking at all of your tests and research and everything that is applicable for that program and then comes up with that strategy. Then with that document you can create tests using Miles, which is the ideation agent. So everything will map back to your strategy. And then Brixton is the newer agent that allows you to use the connector in Claude to more easily map things to the right Airtable base.

Three agents, three handoffs. Dexter reads a client's history and outputs strategic pillars and problem statements. Miles takes that document and turns it into concrete test concepts, usually four to seven per run. Brixton takes the concepts and pushes them into the client's Airtable base, the system of record every strategist actually works from day to day.

The experimentation flywheel: Dexter to Miles to Brixton A three-stage process diagram. Dexter reads a client's tests and research and outputs strategic pillars and problem statements. Miles turns that strategy document into four to seven concrete test concepts. Brixton maps those concepts into the client's Airtable base, infers the schema on its own, and checks for duplicates before pushing. A loop-back arrow shows Brixton's output feeding the next quarter's Dexter run. THE EXPERIMENTATION FLYWHEEL STEP 1 DEXTER Reads a client's tests and research. Outputs strategic pillars and problem statements. STEP 2 MILES Turns the strategy document into concrete test concepts. Usually 4 to 7 per run. STEP 3 BRIXTON Maps concepts into the client's Airtable. Infers the schema. Checks for duplicates before it pushes. Feeds the next quarter's Dexter run speero.com

The room didn't lean in for the ideation, that's table stakes for an AI agent now. It leaned in when I handed Brixton almost nothing and asked it to push six new concepts to Airtable.

It thought quite a bit and actually just had this conversation with Kairi that different bases have things named differently and different things are more important than others. And it actually went out of its way to figure that out on its own without being prompted to do so. It even realized that three of the six tests were duplicates in Airtable already.

Brixton wasn't told the schema, it inferred it. And it caught an overlap a human reviewer, working fast at the end of a long week, might easily have missed.

I'll be honest about why the overlap was probably there, I'd likely run a similar pass on the same client before. That doesn't take anything away from the catch though. An agent that checks its own work before it writes to a shared system is doing something closer to judgment than autocomplete.

Ben Labay, our CEO, was in the room, and his read of the demo went further than the three agents on screen. He wants every client to eventually have one agent, not five separate names to remember:

Ultimately I want every client to have their own agent, not our own project, but their own agent that we onboard and that's connected to everything. And that is instead of having five different agents to remember the names of, like Miles and Brixton and this and that. It's one.
Ben Labay, CEO of Speero

That single sentence turned out to be the whole rest of this year's story.

Why "an agent came up with this" doesn't scare clients

There's a second reason the Dexter to Miles to Brixton chain matters, and I'm not talking about the tech. It's how we talk about the tech once a client is in the room.

A month earlier, on April 14, I called out a pattern I'd noticed across my own accounts:

...Miles is this agent that cross references our current strategic pillars, the problem statements, all of your past test data, and then vets your top five opportunities to consider. That becomes more of a strategic discussion, as opposed to saying these are the exact concepts that we're going to run that I came up with. I just feel like clients may feel more able to say no, or like they're not going to hurt our feelings if they don't approve a test concept, if they believe we individually came up with it.

I flagged this as anecdotal, not a controlled finding. Take it as a working hypothesis, not a law of client psychology, but it's still useful for how you run your own client conversations. Naming the agent turns a rejection of an idea into a rejection of a first draft, not a rejection of a person. That's a small framing shift, but it has a real effect on how comfortable a client feels pushing back, which is exactly the muscle a healthy experimentation program needs them to keep using.

The wild west problem every growing skills library hits

Here's where the story gets less tidy, and where you'll likely recognize your own team if you've shipped more than one or two skills already. Building three well-named agents is the easy part. We spent the spring finding out what happens once a dozen more get built without anyone agreeing on where they live.

Jeff Kellner, US Team and Operations Lead at Speero, said the quiet part out loud in March:

...I'm not afraid to admit I don't know where to find stuff. Where are our ideation agents? And that's a huge barrier of friction for me personally.
Jeff Kellner, US Team and Operations Lead at Speero

By April, the friction had a name and a real cost attached to it. Jonny Longden, Chief Growth Officer at Speero, pushed the team to stop building tools in isolation:

The overarching thing that we came out with was the need to approach this architecture in a top-down way, because it's probably very easy to think of individual applications for things and go off and do them separately, where then you're actually duplicating information or duplicating pieces of work.
Jonny Longden, Chief Growth Officer at Speero

He also raised something sharper than convenience: dependency risk. Every skill built to live entirely inside one vendor's product is a skill that vendor can strand:

The whole system should be transferable, in the sense that it isn't based on any particular tool. In a week Claude could completely change their commercial model and make the whole of Claude completely unviable for us. If everything was inside Claude, you'd be in a problem.
Jonny Longden, Chief Growth Officer at Speero

Jason Lively, PM Lead at Speero, was living the practical version of that same worry. Every skill the team shipped was a plain text file sitting inside Claude's UI. That's close to what Anthropic's own documentation on building agent skills describes as a folder of instructions and reference material. But it had no version history and no way to scope who it applied to:

Right now it's a very laborious process. There's no versioning. There's no version history at all for skills.
Jason Lively, PM Lead at Speero

He named the second cost too, the one that sounds abstract until you've paid it: every skill you connect to a model is context the model has to carry, whether or not the current conversation needs it.

The other issue we've foreseen is the context window bloat. We don't want to expose the whole skill to the model, because that's a huge context window.
Jason Lively, PM Lead at Speero

Vasisth Kumar, Analyst at Speero, put it in the plainest terms of the whole discussion:

If we just connect Claude to everything, and then we say hey, do this, do this, and do that, it kind of confuses the LLM. If I develop a skill for post-test analysis, it gets applied to anyone and everyone, which can confuse Claude.
Vasisth Kumar, Analyst at Speero
Too many names to hold in your head: the agent roster by mid-2026 A grid of Speero's internal AI agents as of August 2026, grouped by function. Ideation and strategy: Dexter, Miles, Brixton. Orchestration: Blast, Rema. Onboarding and PM: Sherpa, Otus. Stats and QA: Testmaster. The uneven column heights show how unevenly the roster grew. TOO MANY NAMES TO HOLD IN YOUR HEAD Speero's internal agent roster, as of August 2026 IDEATION & STRATEGY Dexter Strategy synthesis Miles Test ideation Brixton Airtable mapping ORCHESTRATION Blast Swarms of sub-agents Rema Client data container ONBOARDING & PM Sherpa New-hire onboarding Otus Internal PM agent STATS & QA Testmaster Stats best-practice agent Every column grew at a different pace, and nobody was tracking the total. speero.com

That's four different people, across three months, independently naming the same underlying problem from four different angles:

  • Discoverability.
  • Architecture.
  • Version control.
  • And token cost.

None of them were talking about the same meeting. That's usually the sign a problem is real and not a one-off complaint, and it's worth checking your own library against those four angles before you add a thirteenth skill to it.

Putting your skills library in one place that isn't a chat window

The fix we landed on was a boring, correct engineering decision, not a new agent. It's probably not the fix you want either, if you're hoping for something flashier. Move the skills out of the chat product and into a place with a paper trail.

Jason Lively described the shape of it in April, before it had fully shipped:

We have skills living as MD files inside a repo, instead of packaging them up as a zip or skill file. This would be more programmatic. It could be done through GitHub, and it would be version controlled through there as well. Ideally we'd all be releasing to a staging branch and then merging to main.
Jason Lively, PM Lead at Speero

That's a description of a private version of the same idea Anthropic later formalized publicly as Claude Code plugins, bundles of skills, agents, and connectors distributed through a Git-based directory instead of a chat window's settings panel.

By August, that plan had become policy. Ben Labay described what a real handoff between two people now looks like on a working project:

This whole thing is in a GitHub repository, like all the context files, all the planning files, all of the stuff. So it's not in some my own little private corner in a cloud project. It's all in Git. And then I'm saying, hey, upload that in Git, make sure that's updated in Git. And then Mihkel knows it's there. And then we have the Git actions that are there.
Ben Labay, CEO of Speero

The connective tissue underneath that handoff is the Model Context Protocol, the open standard that lets Claude read from and write to an external system like GitHub in the first place, rather than being limited to whatever context a person pastes into a chat.

Two weeks later, on August 25, the policy became a deletion. Ben announced he was pulling every skill that lived only inside Claude's own organizational settings:

I'm about to delete all of the org skills in Claude. And therefore the only source of truth is going to be in GitHub, where we have that MCP that connects to the GitHub. So we have one source of truth for skills.
Ben Labay, CEO of Speero

That decision closes the loop Jonny Longden opened in April. A skill that only exists as a saved setting inside one product is a skill you can lose the day that product's terms change. A skill that lives as a file in a Git repository, with a commit history and a pull request trail, survives any single vendor's next pricing update.

It also, almost as a side effect, fixes the discoverability problem Jeff Kellner named back in March. You can't grep a chat window's settings panel. You can grep a repository.

From sprawl to single source of truth: March to August 2026 A horizontal timeline with four milestones. March 2026: agent and gem sprawl, no shared categorization. April 2026: governance debate, a GitHub-backed skill plan proposed. May 2026: the Dexter, Miles, Brixton chain demoed live. August 2026: org skills deleted from Claude's UI, GitHub and MCP become the single source of truth, context engineering reframed as level one of harness engineering. FROM SPRAWL TO SINGLE SOURCE OF TRUTH March to August 2026 MARCH 2026 Agent and gem sprawl No shared naming or categorization across the team's agents. APRIL 2026 Governance debate Top-down architecture vs. vendor lock-in risk. GitHub-backed plan proposed. MAY 2026 The chain demoed live Dexter to Miles to Brixton, in front of the room. Brixton catches duplicates. AUGUST 2026 Single source of truth Org skills deleted from Claude's UI. GitHub plus MCP take over. Context engineering becomes level one of harness engineering. speero.com

This is also where the informal agents from the May demo start showing up as something more disciplined. The current Dexter skill in Speero's library isn't a loose prompt anymore.

It opens with a mandatory intake gate: client name, primary KPI, quarterly test velocity, and an explicit choice between three starting modes, before it will touch a single file. It runs a two-pass evidence review, states its file counts out loud, and refuses to classify an ambiguous test result without flagging it for a human.

The MILES skill that replaced the original ideation agent won't generate a single test idea first. It crawls the live page, retrieves test history in two separate passes, and diagnoses the page against six specific lenses before it writes anything.

Then it builds a mandatory duplication check before anything gets pushed to a client's Airtable. That's the exact failure mode Brixton caught by hand back in May, now built into how Speero approaches AI ideation as a required step instead of a lucky catch.

Q-BeRt, the newer skill that assembles client QBR decks, carries its own provenance rule. No number survives into a client deck unless it can be traced back to a live source this quarter. No status gets softened to sound better than the data says it is.

That density is the whole point, and it's also the thing that was about to get a new name.

Context engineering was never the destination

On August 18, Ben Labay stood in front of the agency again, this time with a different message than May. He'd been watching a short video about a term that had started circulating: agent harnesses. And he connected it directly back to the skills the team had spent months building.

When we create a skill, we're doing context engineering, because that skill has a lot of background information. That skill can have reference files that it's referring to. When I use an agent like Sassy or Poly, that's a context engineering problem solved, or applied, if you will. Effectively, where we are now, and where we're going, I think, is really advanced context engineering.
Ben Labay, CEO of Speero

Every one of the fixes in this piece, the GitHub repository, the mandatory intake gates, the phased evidence reviews, is context engineering by that definition. It's giving an agent enough background material, structure, and reference files that it stops guessing and starts working the way Dexter, MILES, and Q-BeRt now do. That's most of what separates a useful skill from a clever prompt, and it's worth applying the same rigor to whatever skills your own team is building.

But Ben's next sentence is the one that reframes everything before it:

Where I think we need to go is harness engineering.
Ben Labay, CEO of Speero

Speero didn't coin that term. Ben was reacting to a short YouTube explainer making the rounds on agent harnesses. The same distinction has since been written up by Martin Fowler and covered on O'Reilly's Radar.

A harness is the surrounding system of tools, checks, and guardrails an agent operates inside. That's distinct from the prompt or the context you hand it. Speero's contribution is applying that distinction to an actual experimentation program instead of a coding assistant, not the vocabulary.

The distinction Ben draws is about the scaffolding, meaning what an agent is allowed to do, what it's connected to, and how its outputs get checked before anything ships. Better prompts and richer files are a different layer, one Speero already solved with context engineering.

A single well-stocked skill still assumes a human is choosing which tool to run and reading the result. A harness is the system around that choice, the thing that governs the whole flywheel from ideation to research to Airtable to the next QBR without a person manually kicking off each step. Ben framed the shift in terms of what the job becomes on the other side of it:

We all need to start thinking about our job shifting to being managers of these agents and sub-agents, like an engineering manager manages engineers. We all need to turn into experimentation managers of experimentation agents.
Ben Labay, CEO of Speero

He went further than a tooling upgrade. He tied it to what kind of company Speero is becoming:

In the end, this is growth engineering. And if we do go this direction, which I think we will, this makes Speero not a growth experimentation agency, but way more of a growth engineering agency.
Ben Labay, CEO of Speero

Jeff Kellner's reaction in the room captured the size of that shift better than any slide could:

It's like the model where agencies would come in and fix your Salesforce, except we're also giving them the Salesforce. It's our software, and we're servicing it.
Jeff Kellner, US Team and Operations Lead at Speero

That identity shift isn't hypothetical. Speero's own homepage already describes the company as a growth engineering agency, not a growth experimentation one.

That brings the story back to where it started. The May 19 demo I ran felt like the finished version of something at the time:

  • three named agents,  
  • a clean handoff,
  • a duplicate catch that impressed the room.

Three months later, that same chain got described as level one of a longer ladder, advanced context engineering, with harness engineering as what we're actually building toward now.

The skills library isn't a single blog post, because it isn't a single thing. It's the record of an agency figuring out, in real time and mostly out loud, the difference between writing a good prompt and building a system that can be trusted to run itself.

This piece is the pillar. Three things come next, roughly in the order you'll hit them yourself:

  1. How you actually govern a skills library once it passes ten contributors.
  2. Why an agent like MILES needing 20 or 30 revisions is a feature of the process, not a warning sign.
  3. And why, past a handful of internal agents, your binding constraint on adoption stops being capability and starts being whether anyone on your team can remember the tool exists.

Did you like this article?

(Your feedback helps us write better!) worst 1 - 10 best

Did the article resonate with you?

What aspects did you enjoy or find lacking?
Were there elements you felt we should've covered?
Thank you.
Oops! Something went wrong while submitting the form.

Related Posts

Who's currently reading The Experimental Revolution?