How-to
How to run growth experiments with AI agents
Run growth experiments with AI agents by joining your data, letting agents hypothesise and ship variants, gating the calls humans should own, and measuring against revenue.
To run growth experiments with AI agents, join your data into one layer, let agents propose and ship variants against a KPI, keep humans on hypothesis selection and ship-to-production, and measure every test against revenue so each result sharpens the next. The agents do the volume; people keep the judgement.
TL;DR in 60 seconds
- Start with joined data. An agent can only reason well across paid, web, lifecycle, and CRM if those sources sit in one view.
- Split the work. Agents draft hypotheses, build variants, and read results; humans approve which tests run, the copy, and what ships.
- Measure against revenue. Grade experiments on margin and conversion, not clicks, so wins are real.
- Run in parallel. Agents can hold dozens of tests across channels at once, which lifts experiment velocity past what a human team can staff.
- Log everything. A structured record of what won, lost, and why is the compounding asset.
The loop, step by step
A growth experiment run by agents is the same scientific loop your team already knows, executed faster and logged properly.
- Hypothesise from the joined data. The agent scans behaviour across the funnel and surfaces where a change is most likely to move a KPI: a drop-off on a pricing page, a dead lifecycle email, a paid segment converting below cost.
- Score and select. Candidates get ranked against effort and expected impact. A human picks which ones run. See prioritising the backlog with ICE, PIE, and RICE.
- Ship the variant. The agent builds the test: new copy, a reworked page section, an audience change. A human approves the copy and the push to production.
- Measure against revenue. Results are read against the outcome that matters, not a vanity metric, using an instrumented funnel.
- Log and sharpen. Win or lose, the result and its reasoning feed the next hypothesis. Nothing is thrown away.
What agents do vs what humans gate
The split is the whole design. Agents are strong at breadth, memory, and tireless execution. Humans are strong at taste, context, and risk.
| Agents own | Humans gate |
|---|---|
| Scanning data for opportunities | Which hypotheses actually run |
| Drafting variants and copy | Copy approval and brand voice |
| Reading and logging results | Ship-to-production |
| Holding many tests in parallel | Strategic direction and spend |
This is not automation for its own sake. It is a guardrail: the cheap, reversible work is delegated, and the decisions with real cost stay with people.
Running in parallel across channels
A single operator tends to run experiments one surface at a time because attention does not divide. Agents do not have that limit. One agent can test a paid audience while another reworks a lifecycle sequence and a third trials a homepage section, all reporting into the same log.
The effect is compounding. In practice we have seen an inbound engine go from zero to roughly 100 leads a month on no paid budget by running many small, measured tests rather than betting on one big campaign. The point is not the number; it is that velocity plus honest measurement beats a few large guesses. More on why in experiments that move margin.
Common mistakes to avoid
- Testing on vanity metrics. If the KPI is not tied to revenue or margin, a "win" can quietly cost you money.
- Skipping the log. Unlogged results mean the system relearns the same lessons.
- Removing the human gate. Letting agents ship unreviewed copy is how you get a fast, confident mistake.
FAQ
Do I need a data warehouse before I start?
No. You need enough of your data joined that an agent can reason across the relevant surfaces. Many teams begin with one funnel step that has clean events and a clear KPI, wire an agent to it, and expand once the loop is proving out. A full warehouse helps later but is not the entry price.
How many experiments can agents actually run at once?
More than a human team can, because agents hold state across many tests without losing attention. The practical cap becomes your traffic and your human approval throughput, not headcount. That is the shift: experiment velocity starts scaling with compute rather than with how many people you can hire.
What stops an agent from shipping something off-brand?
The human gates. Copy approval and ship-to-production stay with people by design, so nothing reaches a customer without a person signing off. Agents propose and build; they do not release unreviewed. The system is fast because the safe steps are delegated, not because oversight is removed.
How do I know an experiment worked?
Measure it against revenue or margin on an instrumented funnel, not against clicks or impressions. A test wins only if it moves the outcome you set at the start. Logging the result and its reasoning turns each answer into an input for the next hypothesis.
Cadence builds and runs this loop inside your own tooling, so the agents, the data layer, and the log stay with you.