Skip to main content
the boring digital co.
BLOG / FRACTIONAL CMO

Why most A/B tests don't tell you what you think they do.

Most A/B tests fail on math, not method. Here's why small businesses misread their test results, and how to run experiments that actually change decisions.

Jack Gamble Jack Gamble, MBA
Co-founder · Marketing, Operations & Project Strategist

Most A/B tests don't tell you what you think they do because they end before the math can support a decision. You change a button, watch the number move, and call it a win. But the number moved for reasons that have nothing to do with your change — a slow week, a holiday, a batch of low-intent traffic, or plain chance. The test looked like proof. It was a coin flip in a suit.

I've watched owner-operators make real budget calls off tests that never earned the right to be trusted. The problem is rarely bad intentions. It's that A/B testing looks simple and is not. Below is what actually goes wrong, and how to run experiments that hold up when you spend money against them.

Why does a "winning" test so often lose in real life?

A winning test loses in real life because it hit significance by accident, not because the change works. Small samples swing hard. If 40 people saw version A and 45 people saw version B, one extra sale can flip which version "wins." You are reading noise and calling it a signal.

Here's the shape of the mistake. A San Diego dental practice tests two versions of its "book a cleaning" page. After a week, version B has a 4.1% conversion rate and version A has 3.2%. That's a 28% lift. The owner ships version B and expects more bookings. Next month, bookings are flat. The lift was real in the sample and imaginary in the world.

The reason is sample size. With small numbers, the range of plausible outcomes is huge. A 4.1% rate off 300 visitors could genuinely be anywhere from 2% to 6% once you account for chance. You didn't measure a difference. You measured a wide guess and picked the friendly end of it.

This is the single most common failure I see. People stop the test when the numbers look good instead of when the numbers become trustworthy. Those are not the same moment.

What is statistical significance, in plain terms?

Statistical significance is a way of asking whether your result is likely real or likely luck. It does not tell you the change is important. It does not tell you the change will hold. It only tells you the difference is unlikely to have happened by chance, given how much data you collected.

Think of it like flipping a coin. Flip it four times, get three heads, and you'd never claim the coin is rigged. Flip it four hundred times and get three hundred heads, now you have a case. Same logic runs your A/B test. Two conversions apart on 80 visitors proves nothing. Two hundred conversions apart on 40,000 visitors is a finding.

Most small business tests never reach the second world. The traffic isn't there. A local law firm might get 1,200 visitors a month to its practice-area page. To detect a real 20% lift with confidence, you often need thousands of visitors per version. At 600 per version, the test would need to run for months — and the market would change underneath it before it finished.

That's the hard truth. Many small businesses do not have enough traffic to run a clean A/B test on most pages. Running one anyway doesn't fail loudly. It fails quietly, by handing you a confident-looking number that is wrong.

What are the mistakes that quietly break a test?

The mistakes that quietly break a test are peeking early, testing too many things at once, and ignoring who the traffic is. Each one produces a number that looks clean and means nothing.

Peeking and stopping early. If you check the results every day and stop the moment version B pulls ahead, you have rigged the test toward false wins. The numbers cross and re-cross by chance all the time. Watch long enough and one version will lead — then you freeze the frame at the flattering moment. Decide your sample size and end date before you start. Then don't look until you get there.

Testing too many changes at once. You redesign the headline, the button color, the form length, and the hero image, then version B wins. What won? You have no idea. You can't repeat it, can't explain it, and can't apply the lesson anywhere else. One variable per test. Boring, slow, and the only version that teaches you anything.

Ignoring who showed up. Traffic is not uniform. A version that ran mostly during a paid campaign got different visitors than a version that ran mostly on organic Tuesday traffic. If your split wasn't truly random and simultaneous, you compared audiences, not pages. Run both versions at the same time, to the same traffic sources, split at random.

Measuring the wrong thing. Clicks are not bookings. Bookings are not revenue. A button that gets more clicks but fewer completed forms is a loss dressed as a win. Tie the test to the outcome that pays you. If you can't tie it to revenue, you can't celebrate it.

When should a small business run an A/B test at all?

A small business should run an A/B test only when it has enough traffic to reach a clear answer in a reasonable window. If a page can't produce a few thousand visitors and a few hundred conversions per version inside a month or two, an A/B test is the wrong tool. You will spend weeks and end with a guess.

The good news is that A/B testing is not the only way to learn. Most small businesses get more from methods that don't need huge numbers.

Before-and-after with a long enough window. Change one thing, measure the same period the month before and the month after, and account for seasonality. Not as clean as a controlled test, but honest about its limits — and often the only realistic option at small scale.

Watching real people use the page. Five session recordings will show you where visitors get stuck. That's qualitative, not statistical, but it points at problems you can fix without needing to prove a 0.4% lift.

Fixing obvious breakage first. A form that fails on mobile, a page that loads in six seconds, a phone number that isn't clickable — these aren't test candidates. They're repairs. You don't A/B test whether the front door should open.

The order matters. Foundations first. Fix what's clearly broken, watch how people actually behave, and save formal A/B testing for the few high-traffic pages where the numbers can carry the weight. This sequencing is part of how we think about growth experiments as a Fractional CMO — you run the test the business can actually afford to trust, not the test that sounds impressive in a meeting.

How do you run an experiment that actually changes a decision?

You run an experiment that changes a decision by deciding, before you start, what result would make you act. If no possible outcome would change what you do next, don't run the test. Write down the decision first. The experiment exists to serve it.

Here is the checklist I use with owners:

  1. Name the decision. "If B wins, we roll it out to all four service pages." If you can't finish that sentence, stop.
  2. State the minimum effect worth acting on. A 1% lift on a page that gets 200 visitors a month isn't worth the work. Decide the smallest change that would matter to revenue.
  3. Calculate the sample size before you start. Free calculators do this in a minute. If the number is unreachable, pick a different method now, not after six wasted weeks.
  4. Change one variable. Just one.
  5. Set the end date and don't peek. Let it run to full sample. No early calls.
  6. Check the result against the decision. Did it clear the threshold you set? Act. Did it not? That's also an answer. A test that says "no difference" saved you from shipping a change that would have done nothing.

This discipline is unglamorous. It's also the difference between experiments that compound into real knowledge and experiments that produce a folder of contradictory "wins" nobody can explain. When we worked with McShanes Solicitors, the gains came from clear positioning and fixing what the numbers pointed to — not from chasing marginal lifts on pages that couldn't support a test.

Most owners don't need a testing culture. They need a small number of good decisions made on evidence they can trust. That's a different, smaller, more honest goal.

What we won't tell you

We won't tell you A/B testing is the answer to slow growth. For most small businesses it isn't, because the traffic isn't there to make it work. Selling you a testing program you can't feed would be easy money and bad advice. The bigger levers are usually upstream — who you're talking to, what you say, and whether the basics work. Testing sits at the far end of that list, not the front.

If you're weighing whether to bring in senior help to run this well, it's worth reading when you actually need a Fractional CMO and when you don't before you commit. And if you're choosing between a strategist and an agency to run your experiments, the difference that matters is worth ten minutes. The wrong structure produces plenty of tests and very few decisions.

Run fewer experiments. Run them properly. Believe the ones that earn it.

— FAQs

Things readers usually ask.

How much traffic do I need to run a reliable A/B test?
It depends on your current conversion rate and the size of the lift you want to detect, but most tests need several thousand visitors and a few hundred conversions per version to be trustworthy. A free sample-size calculator will give you the exact number before you start — if it's unreachable in a month or two, pick a different method.
Why did my winning A/B test not improve results after I shipped it?
Because the win was likely chance, not a real effect. Small samples swing hard, and stopping a test the moment one version pulls ahead locks in noise that disappears once the change goes live to everyone.
Can I test more than one change at a time to save time?
You can, but you won't know which change caused the result, so you can't repeat it or apply the lesson elsewhere. Test one variable at a time — it's slower, but it's the only version that teaches you anything reliable.
What should I do instead of A/B testing if I have low traffic?
Fix obvious breakage first, watch real session recordings to find where visitors get stuck, and use honest before-and-after comparisons over a long enough window. These methods don't need large numbers and usually surface bigger problems than a fractional-percent test would.
How do I know if an experiment is worth running at all?
Write down the decision the result would change before you start. If no possible outcome would change what you do next, the test isn't worth running.
— READ NEXT
— GET IN TOUCH

Want us to look at your site?

A 20-minute call. No pitch. We'll tell you what we'd fix first.

CONTACT US →