Email subject line tests: useful exercise or expensive theatre?
Subject line A/B tests can be useful or a waste of time. The difference is your list size, your sample math, and whether you act on the result.
Subject line tests are useful when your list is large enough to produce a real signal, and expensive theatre when it is not. Most small businesses fall in the second group. They run a test on a list of 900 people, watch a 3% open-rate gap, declare a winner, and change nothing about how they write next week. That is not an experiment. That is a coin flip you paid attention to.
I have run these tests inside marketing teams and inside my own campaigns for years. The test itself is cheap. The trap is the false confidence it creates. Below is how to tell whether your subject line testing is teaching you something or just giving you a chart to look at.
Why do most subject line tests tell you nothing?
Most subject line tests tell you nothing because the sample is too small to separate a real effect from random noise. Email open rates bounce around on their own. Two versions of the same subject line, sent to two random halves of the same list, will rarely produce the exact same open rate. The gap you see might be the subject line. It might be the time of day. It might be luck.
Here is the math that matters. To detect a small difference between two subject lines with any confidence, you need thousands of recipients per version — often 5,000 or more each, depending on how big a lift you are trying to spot. A firm with a list of 1,200 does not have that. When they split the list in half and send version A to 600 and version B to 600, the result is close to meaningless. A 4-point open-rate gap on 600 people can flip the next time you run the same test.
Apple's Mail Privacy Protection made this worse. Since 2021, a large share of "opens" are machine-generated pre-fetches, not humans reading. Your open rate is now a blurry number sitting on top of an already-noisy metric. Testing a blurry number with a small sample is how you end up confident and wrong.
The honest read for most small businesses: if your list is under a few thousand active subscribers, a single subject line test cannot tell you which line is better. It can only tell you which line got more opens that one time.
When is a subject line test actually worth running?
A subject line test is worth running when your list is large enough to reach significance, when the two options are genuinely different, and when you plan to change your writing based on what you learn. All three have to be true. One out of three is theater.
List size is the gate. If you send to 20,000 people a week, you have room to split off a test group and still get a signal. If you send to 800, you do not. There is no clever workaround. Small lists produce small samples, and small samples produce noise.
The second condition is contrast. Testing "May newsletter" against "Your May newsletter" is a waste even on a big list. The lines are nearly identical, so the true difference is tiny, so you would need an enormous sample to see it. Test ideas that are actually different. A question against a statement. A specific number against a vague promise. Curiosity against clarity. If a stranger could not tell the two lines apart in intent, do not test them.
The third condition is the one people skip: you have to act on the result. A test that changes nothing is a diary entry. If you learn that specific numbers beat vague promises, that lesson should show up in your next ten subject lines, your landing page headlines, and your ad copy. The value is not the single winning email. The value is the pattern you carry forward.
What should small businesses test instead?
Small businesses should test the things that move money, not the things that are easy to measure. Open rate is easy to measure. Revenue is what you actually care about. Those are not the same, and chasing the easy one is how firms spend a year optimizing subject lines while their booking flow leaks clients.
Start with the click. A subject line's job is to get the email opened. But the email's job is to get a click to your site or a reply to your inbox. If you must test, test the whole chain — subject line, body, and call to action together — against a version built on a different idea. Measure clicks and replies, not just opens. Those are behaviors, not machine pre-fetches.
Better still, spend your testing energy upstream of email entirely. For a service business, the biggest wins usually sit in three places. The first is who finds you in search before they ever join a list. The second is what your service pages say when they land. The third is whether your intake process turns an inquiry into a booked call. A 2-point open-rate lift on a small list is a rounding error next to fixing a contact form that half your visitors abandon.
This is the sequencing question a good Fractional CMO works through before touching email at all. You do not run growth experiments in the order that feels busy. You run them in the order of expected return. Email subject lines are near the bottom of that list for most small firms, not because they never matter, but because the leverage is small and the sample is smaller.
How do you run a test that produces a real answer?
You run a test that produces a real answer by deciding the rules before you send, sizing the sample honestly, and refusing to peek and stop early. Most tests fail on process, not on idea. Here is a sequence that holds up.
- Write down the question. Not "which subject line is better" but "does a specific number beat a vague benefit for this audience." A vague question produces a vague lesson.
- Pick one variable. Change the subject line and only the subject line. Same send time, same segment, same body. If you change two things, you cannot know which one moved the number.
- Size the sample before you send. Use a free significance calculator. Enter your baseline open rate and the smallest lift you would care about. If the calculator says you need 6,000 per version and you have 1,500 total, stop. You cannot answer this question. That is a real answer too.
- Randomize the split. Your email tool does this. Do not hand-pick groups.
- Set the finish line in advance. Decide the sample size and stick to it. Do not check the result at hour two, see a winner, and declare victory. Early peeking is how noise gets crowned.
- Measure clicks and replies, not only opens. Opens are polluted. Behavior is not.
- Write down what you learned in one sentence. "Numbers beat vague benefits by a real margin" or "no detectable difference, our list is too small to test this." Both are useful. Both should shape what you do next.
Run this a few times and you will notice something. The disciplined version is slower and less exciting than the dashboard makes it look. That is the point. Boring by design beats fast and wrong.
What does this look like for a real firm?
For most professional-services firms, the answer is to stop testing subject lines and start fixing the funnel around them. A law firm sending a monthly update to 2,000 past clients does not have a subject line problem. It has a follow-up problem, a search-visibility problem, or an intake problem — and those are where the money is.
When we worked with McShanes Solicitors, the gains came from getting the right people to find the firm and turning those visitors into inquiries. Not from squeezing a few more opens out of a newsletter. That is the pattern almost every time. The clients already looking for you are worth more than the clients you are trying to re-open an email.
Think about the math in dollars. Say a San Diego family-law firm has a list of 1,800. A subject line test that "wins" by 3 points on opens might mean 54 more people opened one email. If your open-to-inquiry rate is 1%, that is half a person. Now compare that to a single new client from search worth $6,000 in fees. The subject line test is not wrong. It is just small. Time spent there is time not spent on the thing that pays.
This is why the sequencing conversation matters more than any single tactic. If you are weighing whether a senior operator should own that sequencing, it helps to read when you actually need a Fractional CMO and when you don't before you spend money either way. And if you are choosing between an outside strategist and a delivery shop, the difference between a Fractional CMO and an agency decides who owns the order of your experiments in the first place.
Where this breaks down
Subject line testing is not useless. If you run a real ecommerce list of 50,000, test away — the sample is there and the lift compounds. This post is aimed at the small firm with a small list and a big to-do list. For you, subject line testing is usually the wrong first experiment, not a forbidden one.
What we will not tell you is that a testing tool will fix a growth problem. A dashboard that reports open rates to two decimal places on a list of 900 is precision without accuracy. It looks like rigor. It is theater. The useful version of this work is slower, less shiny, and tied to revenue. Search first. Funnel aware. Foundation always.
Things readers usually ask.
- How big does my email list need to be to test subject lines?
- You generally need several thousand recipients per version to detect a small difference with confidence — often 5,000 or more each. Lists under a few thousand active subscribers rarely produce a reliable signal from a single subject line test.
- Are email open rates still reliable after Apple's Mail Privacy Protection?
- Open rates are now polluted by machine-generated pre-fetches, so a large share of reported opens are not human reads. Measure clicks and replies, which reflect real behavior, instead of relying on opens alone.
- What should a small business test before subject lines?
- Test the things that move money — search visibility, service page copy, and your intake or booking flow. Those have far more leverage than a small open-rate lift on a small list.
- How do I know if my subject line test result is real?
- Decide your sample size before sending, use a significance calculator, and do not stop early when you spot a winner. If the calculator says you need more recipients than your list holds, the honest answer is that you cannot test that question yet.
- Is subject line testing ever worth it for a small firm?
- It can be, but only when your list is large enough, the two lines are genuinely different, and you plan to apply the lesson to future writing. For most small firms, the sample is too small and the payoff too low to make it a priority.
Want us to look at your site?
A 20-minute call. No pitch. We'll tell you what we'd fix first.
CONTACT US →