Email A/B Testing: What to Test and When to Trust the Result

M
MailGraf
Sep 29, 2026

Email A/B testing means sending two versions of the same email campaign to two halves of your list, then checking which one gets more opens or clicks. Instead of arguing about which subject line is better, you let your own subscribers decide by what they actually do.

The idea is simple. Most tests go wrong in one of two ways: changing several things at once, or deciding with too few people. The sections below cover what to test, how many people you need, how long to wait and how to read the result.

What is email A/B testing?

In an A/B test you change one thing and keep everything else the same. The version you leave alone is Version A. The version with the change is Version B.

Picture a homeware shop announcing its autumn range. It can't decide between two subject lines:

  • Version A: Our autumn range has arrived
  • Version B: 12 new pieces to get your home ready for autumn

The list has 4,000 people. 2,000 get Version A, 2,000 get Version B, and both go out at the same time. The next day the report shows that 24% of the A group opened the email, against 28% of the B group. This audience responded better to the line that told them exactly what they would find inside.

You will also hear it called split testing. Websites and ads use the same method. The advantage with email is speed: both versions go to the same list on the same day, and results start arriving within hours.

One common mix-up is worth clearing up early. Sending different emails to different groups is not an A/B test. Sending a store opening to London subscribers and an online offer to everyone else is segmentation. In an A/B test the two groups are alike and picked at random. The only difference is the version they receive.

Why A/B testing works for email

Much of the time spent on email marketing goes on decisions. Should the subject line be short or long? Should the button say "Shop now" or "See the range"? Everyone has an opinion, and the most confident voice in the room often wins.

A/B testing shortens that debate. The decision rests on what the people receiving your email did, not on anyone's hunch.

The real value builds up over time. Each test teaches you one small thing about your audience: whether a price in the subject line lifts opens, or whether shorter emails get more clicks. After a few months of regular testing you have a set of rules that belong to your own subscribers.

That is why testing works best as a habit in your email marketing strategy rather than an occasional experiment. Start with the emails you send most often, such as your newsletter, a weekly offer or new product announcements. They give you results the fastest.

What can you test?

Everything you might test falls into two groups, and the group decides how you read the result. First, two terms:

  • Open rate: the share of people who received the email and opened it.
  • Click rate: the share of people who received the email and clicked a link in it.
What you changedWhat you look atExample
Subject lineOpen rateA question against a plain statement
Preview textOpen rateNaming the offer against building curiosity
Headline and body copyClick rateShort and direct against a short story
Button textClick rate"Shop now" against "See the range"
ImageClick rateA product on white against the product in a real room
Email lengthClick rate3 products against 8 products

The logic is straightforward. The subject line and preview text decide whether someone opens the email. They don't change what happens inside it. So a content test is judged on clicks, not opens.

Subject line testing

Subject line testing is the most popular place to start, and for good reason: it is quick to set up and every send gives you a result.

Test the subject line and the preview text as a pair. The preview text is the short line shown next to or under the subject in the inbox, and people read the two together before deciding to open. If you change the subject line, write a preview text that fits it. For ideas worth testing, look at emails you have sent before and pick the two styles your team argues about most.

Three ideas for your first test

  • Specific or general? "New arrivals are here" against "8 new pieces for your kitchen". You find out whether people open more readily when they know what is inside.
  • Offer in the subject line or inside? "20% off everything this weekend" against "The autumn edit is here". This tells you whether your audience prefers the offer up front or a little curiosity.
  • One button or one per product? A content test. A single large "See the range" button against a "View" button under each product. You judge it on click rate.

How many people do you need?

This is the question most guides skip. The honest answer is that it depends on how big the difference between your two versions is. A small difference needs a lot of people to prove. A large one needs fewer.

Think about tossing a coin. Getting 6 heads in 10 tosses wouldn't surprise you; that happens by chance. Getting 600 heads in 1,000 tosses would make you suspect the coin. Email tests work the same way. A difference in a small group may be luck. The same difference in a large group is much harder to put down to chance.

The table below keeps the result the same on every row: 24% of the A group opened, against 28% of the B group, a gap of 4 points. The only thing that changes is how many people received the email. The last column shows the sentence the MailGraf report displays in each case.

People who received itPer versionWhat the report says
300150Not enough data yet to call a winner.
500250The difference is small. A larger audience would make it clearer.
1,000500The difference is small. A larger audience would make it clearer.
2,0001,000Enough data to trust this result.
5,0002,500Enough data to trust this result.

These numbers are people who actually received the email. Bounced addresses, the ones that couldn't be delivered, don't count.

If your list is small, don't chase small differences. "Our autumn range" against "Our autumn range has arrived" will never produce a gap you can measure. Test ideas that are genuinely different instead, such as a subject line built around a discount against one built around a new product.

To put a number on it: with 150 people per version, the gap needs to be around 12 points before the result can be trusted. Gaps that large only come from big changes. Below 100 people per version, results are mostly luck, and MailGraf won't name a winner.

How to run an email A/B test, step by step

Whatever tool you use, a good test follows the same six steps.

  1. Write down a question. One sentence is enough: "Does putting the price in the subject line lift opens?" Note what you expect to happen.
  2. Change one thing. Version A is what you usually send. Version B is the new idea.
  3. Split your list at random. If Version A goes to long-standing subscribers and Version B to new ones, the audience makes the difference, not the subject line. A random split keeps the two groups alike.
  4. Send both versions at the same time. A morning email and an evening email don't compete on equal terms. Sending together stops the send time from skewing the result.
  5. Wait. Don't judge on the first hour or two. People read email at different times of day and early numbers swing a lot. Give it at least a full day.
  6. Record the result and use it. Keep a simple log: date, question, Version A, Version B, result. Use the winning idea in your next send.

Before you send an A/B test

  • Have I changed only one thing?
  • Do I know which number I'll judge it on: opens or clicks?
  • Will each version reach at least a few hundred people?
  • Have I sent both versions to myself as test emails?
  • Do the links and the unsubscribe link work in both versions?

How to read the result

The report compares two rates. The gap is usually given in points: between 24% and 28% there are 4 points. You will see one of three outcomes.

  • One version is clearly ahead. Make the winning idea your new default. In the next test, put a fresh idea up against it.
  • The gap is small. Treat it as a hint, not a rule. Run the same idea once or twice more.
  • There is no difference. That is a result too. It tells you this change doesn't matter much to your audience, so spend your next test on something else.

It helps to look at opens and clicks together. If a subject line lifts opens but clicks fall, it may have raised curiosity that the email didn't satisfy. Our email marketing KPI guide explains what each of these numbers can and can't tell you.

Why open rates need a pinch of salt

When an Apple Mail user turns on Mail Privacy Protection, the sender can't see whether they opened the email (Apple). Opens from these people aren't reliable; an email can look opened even if nobody read it.

That is not a small group. In Litmus's July 2026 data, 62% of tracked email opens came from Apple, and Litmus does not treat privacy-protected opens as reliable (Litmus Email Client Market Share).

A/B testing in MailGraf

In MailGraf, A/B testing sits inside the email campaign wizard, so there is no separate tool to open.

  1. In the campaign wizard, click the settings icon at the top right and turn on Run an A/B test.
  2. Choose Subject line or Content. A subject line test takes two subject lines and preview texts. A content test takes two designs.
  3. For a content test, click Copy from A to start from the same design and edit only the part you want to test. Show comparison opens both designs side by side.
  4. When you send a test email, choose Both under Send as to receive both versions.
  5. On the Audience step, the A/B split card shows how many people get each version. Then send.

Both versions go out at the same time and every recipient gets only one. The split is balanced across mailbox providers, so Gmail, Outlook and Apple Mail users are shared evenly between the two versions. One version can't end up with more Apple users by chance and skew the result.

The result is on the A/B test tab of the report. The card at the top shows which version is ahead and by how many points. The line underneath tells you whether you can trust it. It only says "Enough data to trust this result." when two conditions are both met: each version has reached at least 100 people, and the gap is too clear to be explained by chance.

MailGraf doesn't hold back part of the list and send the winner later; both versions go out together from the start. The upside is no waiting: your whole email campaign leaves at the time you planned. The trade-off is that half your list receives the weaker version. You apply what you learn in your next send.

For a screen-by-screen walkthrough, see the help article Run an A/B test; for an overview of the feature, see the A/B testing page. A/B testing works alongside lists, segments and reports in the same place; you can see the rest of the toolkit on our features page.

Common mistakes

  • Changing several things at once. You get a winner but no idea why.
  • Deciding in the first few hours. The people who open in the first hour or two don't represent your whole list. Wait at least a day.
  • Chasing small gaps on a small list. Pitting two similar subject lines against each other on 300 people usually ends with "not enough data". On a small list, test big ideas.
  • Turning one result into a permanent rule. A winner once isn't a winner forever. Seasons, offers and lists change, so test important findings more than once.
  • Not writing the result down. An unrecorded test means having the same debate again in three months.
  • Looking at opens only. If opens rise and clicks fall, the "winning" version may not have won at all.

FAQ

How many people do I need for an email A/B test?

There is no single number; it depends on how big the gap between the versions is. Below 100 people per version, results are mostly luck. With 1,000 people per version, a 4-point gap becomes trustworthy. On a small list, only large gaps give a reliable answer.

How long should I wait before reading the result?

Give it at least a full day. People read email at different times and early numbers swing a lot. For emails sent at the weekend, a little longer is sensible.

Does half my list get the weaker email?

When both versions are sent at the same time, yes: half the list gets the version that turns out weaker. In return there is no waiting, and the whole email campaign leaves at the time you planned. You use what you learn in the next send.

Can I test the subject line and the content at the same time?

Not in the same test. If both change, you can't tell which one made the difference. Test the subject line first, then the content in another send.

Why does my open rate look so high?

Apple Mail Privacy Protection can make an email look opened even when nobody read it, which pushes open rates up. In an A/B test the effect is spread across both versions, but checking clicks is a good way to confirm the result.

Does the winning version always win?

No. A test tells you what worked better on that day, with that list and that offer. Testing an important finding a few more times across different email campaigns gives you a sturdier rule.

Can I A/B test automated emails?

In MailGraf, A/B tests are available on email campaigns, not on automations. If you want to try a new idea for an automated email such as a welcome series, test it in an email campaign first.

Originally published: Sep 29, 2026

MailGraf

Professional email marketing platform.

Don't miss out

Get the latest email marketing tips and exclusive updates.

ISO CertifiedGDPR CompliantCSA Certified

MailGraf is a trading name of MailGraf Digital Ltd, registered in England and Wales, No. 13282175. ICO ZB250899.