Growth experiments: how to design one that proves something.
A growth experiment proves something only when it has one clear hypothesis, one metric, and a stopping rule. Here's how to design one that gives you a real answer.
A growth experiment proves something when it starts with a single hypothesis, measures one number, and defines in advance what result would change your mind. Most of what gets called an experiment is just a change made and watched. That is not an experiment. That is a hope with a dashboard.
The difference matters because you are spending real money and real time. When you run a test that cannot fail, you learn nothing. When you run a test that can fail, you learn something either way. This post is about the second kind — how to design one, run it, and read the result without lying to yourself.
What makes something an experiment instead of a change?
An experiment has a hypothesis you could be wrong about. A change does not. That is the whole distinction, and it is the one most business owners skip.
Say you swap the headline on your services page. If you tell yourself "the new one is better," you have made a change. You will look at next month's numbers, they will be up or down for a dozen reasons, and you will credit or blame the headline. That is a story, not a finding.
An experiment says: "If the new headline is clearer, then the form-fill rate on this page will rise by at least 15% over the next 300 visitors." Now you have something that can be true or false. You named the mechanism (clearer copy), the metric (form-fill rate), the size that would matter (15%), and the sample that would settle it (300 visitors). If the number does not move, the idea was wrong. That is a good outcome. You just saved yourself from rolling out clearer copy everywhere for no reason.
The test of a real experiment is simple. Ask yourself before you start: what result would make me abandon this idea? If you cannot answer, you are not running an experiment. You are decorating a decision you already made.
How do you write a hypothesis that can actually fail?
A good hypothesis names the cause, the effect, and the number — in one sentence you would be willing to lose. The format is boring on purpose: "If we do X, then Y will change by Z, because of some mechanism we can explain."
Here is a weak one: "Adding client reviews will build trust." It cannot fail. Trust is not measured, X is vague, and there is no number. You could add reviews, watch anything happen, and declare victory.
Here is a stronger version: "If we add three named client reviews to the fees page, then the rate of visitors who reach the contact form will rise by at least 10%, because social proof reduces hesitation about price." Now you can be wrong. The rate might not move. It might drop, if the reviews clutter the page. Either way you learn.
Three parts every hypothesis needs:
- The change, stated precisely. Not "improve the page." Say what you will add, remove, or move.
- The metric, chosen before you start. One number. If you pick the metric after you see the data, you will pick the one that looks good.
- The threshold that matters. A 2% lift on 40 visitors is noise. Decide what size of effect would change what you do next. Below that, treat it as no effect.
Write the hypothesis down and date it. The act of writing it forces the honesty. Half-formed ideas fall apart the moment you try to put a number on them, and that is the point.
Which metric should you measure — and which one lies to you?
Measure the metric closest to money that the experiment could plausibly move. Vanity metrics — traffic, impressions, followers — sit too far from revenue to prove anything about a small change to one page.
Most of the time the honest metric is a conversion rate, not a count. Counts go up when traffic goes up, and traffic goes up for reasons that have nothing to do with your test. If your homepage got 40% more form fills last month, the cause could be your new copy, or a competitor closing, or a seasonal swing, or a single referral you never noticed. A rate — form fills divided by visitors — strips most of that noise out.
Pick one primary metric. You can watch others, but you commit to one before you start, and that is the one that decides the call. When you let yourself grade the experiment on any of five metrics, one of them will always look good by chance, and you will always find a way to declare the test a success. This is the most common way owners fool themselves. If we cannot tie a result to revenue, we do not celebrate it, and a single committed metric is how you keep that promise to yourself.
For a service business, the metrics that usually matter are: qualified enquiries, booked consultations, and closed clients. Everything above those in the funnel is a leading indicator — useful, but not proof on its own.
How long do you run it, and how do you know when to stop?
Decide the sample size and the end date before you launch, and do not stop early because the result looks good. Stopping early is the most expensive mistake in this whole discipline, because early results swing wildly and then settle.
You do not need to be a statistician. You need a stopping rule. The simplest one is a minimum sample: "I will not read the result until at least 300 visitors have seen each version." For a low-traffic page, that might take six weeks. For a busy one, three days. The point is that you named the number in advance, so you cannot quit the moment the chart tips your way.
Why this matters: in the first 50 visitors, a variant can look 40% better and end up 5% worse. Random noise looks like a signal when the sample is small. If you check daily and stop the instant one version leads, you will "prove" things that are not true, over and over, and build your whole strategy on coin flips.
A workable process for a small business:
- Set the primary metric and the threshold that would matter.
- Estimate how many visitors or leads you need. If you have no idea, 300 conversions events per variant is a rough floor for a rate-based test; fewer than 100 and you are mostly reading noise.
- Set an end date as a backstop, so a test that never reaches sample size does not run forever.
- Do not look at partial results as if they were final.
- When the sample is reached, read the primary metric and only the primary metric.
Low-traffic businesses hit a wall here. If your site gets 400 visitors a month, a clean A/B test on a single page could take a full quarter to settle. That is real, and it changes what you should test. On thin traffic, test bigger swings — a whole new page structure, a different offer — where the effect is large enough to see fast. Save the small tweaks for pages that get enough traffic to measure them.
What should you do with the result once you have it?
Write down what you predicted, what happened, and what you will do next — before you move on. An experiment you do not record is an experiment you will run again by accident in eight months.
There are three outcomes, and all three are useful.
The metric moved past your threshold. Ship the change. Then ask whether the same mechanism applies elsewhere. If clearer copy lifted conversions on the fees page, the lesson is not "that headline works" — it is "our pages are unclear about price." That insight is worth more than the single win.
The metric did not move. The idea was wrong, or the effect was too small to matter. Either way you stop investing in it. This is a result, not a failure. A test that kills a bad idea before you scale it across the site saves more money than most wins make.
The result was ambiguous — a small move inside the noise. Treat this as no effect. Do not squint at it until it becomes a yes. If the change was cheap and harmless, keep it and move on. If it cost something, drop it. Ambiguous is not a reason to run it three more times hoping for clarity.
The discipline compounds. Ten experiments a year, each with a written prediction and an honest read, teaches you more about your own customers than a decade of untested changes. When we ran positioning tests for McShanes Solicitors, the value was not any single winning page — it was the pattern that emerged across tests about what their clients actually needed to hear before they picked up the phone.
This is the kind of work a Fractional CMO owns — setting the hypothesis, guarding the metric, and refusing to celebrate a result that will not hold. It is unglamorous and it is where most of the real gains hide. If you are weighing whether you need that role at all, we wrote about when a Fractional CMO earns their keep and when they don't, and about how the role differs from hiring an agency.
Where this breaks down
Growth experiments do not work on tiny traffic with tiny effects. If your page gets 200 visitors a month and you are testing a button color, you will wait a year for an answer that never arrives. On thin traffic, test big swings or do not test at all — use judgment and your foundation instead. Experiments also cannot fix a broken offer. If the thing you sell does not match what people want, no headline test will save it, and running twenty of them just delays the harder conversation. Test the offer before you optimize the page.
Things readers usually ask.
- How much traffic do I need to run a growth experiment?
- As a rough floor, aim for at least 300 conversion events per variant before you read a rate-based test; below 100 you are mostly measuring noise. If your traffic is thin, test large changes where the effect is big enough to see quickly rather than small tweaks.
- What is the difference between an experiment and just making a change?
- An experiment has a written hypothesis you could be wrong about — a named change, one metric, and a threshold decided in advance. A change is something you do and then watch, crediting or blaming it based on numbers that moved for many reasons.
- Why shouldn't I stop an experiment early when it looks like it's winning?
- Early results swing wildly and often reverse once more data arrives; a variant that looks 40% better after 50 visitors can end up worse. Set your sample size before you start and do not read the result until you reach it.
- What metric should a growth experiment measure?
- Measure the metric closest to money that your change could plausibly move — usually a conversion rate like qualified enquiries or booked consultations, not traffic or impressions. Commit to one primary metric before you start so you cannot grade the test on whichever number happens to look good.
- What if the result comes back ambiguous?
- Treat a small move inside the noise as no effect and do not run the test repeatedly hoping for clarity. If the change was cheap and harmless, keep it and move on; if it cost something, drop it.
Want us to look at your site?
A 20-minute call. No pitch. We'll tell you what we'd fix first.
CONTACT US →