Simple steps to use A/B testing for optimising email campaigns
Email marketing is a discipline where small, data-driven adjustments can compound into dramatic improvements in engagement and revenue. A/B testing (also known as split testing) is the most reliable method to discover what truly resonates with your audience - replacing guesswork with evidence. By systematically testing different elements of your emails and measuring the results, you can make confident, data-backed decisions that steadily enhance the performance of every campaign you send.
According to a 2024 Invesp study, companies that run A/B tests on their email campaigns see an average 37% improvement in click-through rates and a 20% increase in revenue per email compared to untested campaigns. Yet surprisingly, Litmus reports that only 39% of brands regularly A/B test their emails - meaning there is a significant competitive advantage waiting for those who commit to a testing culture.
In this comprehensive guide, we will walk you through every step of the A/B testing process - from understanding the fundamentals to advanced multivariate techniques - and show you how platforms like MailCraft make testing accessible, rigorous, and actionable.
1. Master the fundamentals of email A/B testing
Before launching your first test, it is essential to understand the core principles that make A/B testing valid and useful:
- What A/B testing is: A method where two versions of an email element (Version A and Version B) are sent to randomly selected, equally sized subsets of your audience. The version that performs better against a predefined metric is declared the winner and can be deployed to the rest of your list.
- The purpose: To identify which specific changes yield measurable improvements in metrics like open rates, click-through rates, conversion rates, or revenue per email.
- The single-variable rule: For a test to produce clear, actionable results, you should change only one element at a time. If you change both the subject line and the CTA button colour simultaneously, you cannot determine which change caused the difference in performance.
- Statistical significance: A result is only meaningful if it is statistically significant - meaning the observed difference is unlikely to have occurred by chance. Most practitioners aim for a 95% confidence level, which means there is only a 5% probability that the result is random noise.
A common pitfall is ending tests too early. VWO research shows that 70% of marketers check their A/B test results within the first 48 hours - often before enough data has been collected for statistical significance. Patience is a prerequisite for reliable insights.
Understanding these fundamentals protects you from common mistakes and ensures that every test you run produces genuinely useful insights.
2. Define clear, measurable objectives before testing anything
Every A/B test should start with a specific hypothesis and a clear success metric. Testing without a defined objective is like navigating without a compass - you may learn something, but you won’t know if it matters.
- Set specific goals: Determine whether you are trying to increase open rates, boost click-through rates, improve conversion rates, reduce unsubscribes, or maximise revenue per email. Each goal will influence which email element you choose to test.
- Formulate a hypothesis: Frame your test as a prediction. For example: “Adding the recipient’s first name to the subject line will increase open rates by at least 10% compared to a generic subject line.” A clear hypothesis gives you a benchmark against which to evaluate results.
- Choose your primary metric: While it is valuable to track multiple metrics, designate one as your primary success metric to avoid confusion. If you are testing subject lines, the primary metric should be open rate. If you are testing CTA button design, focus on click-through rate.
- Align with business outcomes: Always connect your testing programme to larger business goals. Improving open rates is valuable, but ultimately, you want those opens to drive clicks, conversions, and revenue. Prioritise tests that have the highest potential to impact the metrics that matter most to your bottom line.
Having clear objectives ensures that your testing efforts are purposeful, prioritised, and directly connected to your overall marketing strategy.
3. Choose the right email elements to test
Not all email elements are created equal in terms of their potential impact. Select the components that are most likely to influence your primary metric, and always test one variable at a time:
- Subject lines: This is the single most impactful element for open rates. Test different approaches: question vs. statement, short vs. long, personalised vs. generic, with emoji vs. without, urgency-driven vs. curiosity-driven. Campaign Monitor research found that subject lines with 6-10 words deliver the highest open rates.
- Preheader text: The preview text that appears next to (or below) the subject line in most email clients. Despite being the second thing recipients see, many marketers leave this as default filler text. Test whether a compelling preheader boosts open rates.
- Email body copy: Test variations in tone (formal vs. conversational), length (concise vs. detailed), structure (paragraphs vs. bullet points), and storytelling approach (problem-solution vs. benefit-led). Shorter emails often outperform longer ones for promotional campaigns, but educational content may benefit from depth.
- Call-to-action (CTA) buttons: Test the CTA text (“Shop Now” vs. “Get My Discount” vs. “See the Collection”), button colour, button size, and placement (above the fold vs. after the content). WordStream data indicates that personalised CTAs convert 202% better than generic ones.
- Images and visual elements: Compare emails with a hero image versus text-only, product photography versus lifestyle imagery, or single image versus image gallery. For some audiences, a clean text-based design outperforms image-heavy layouts.
- Send times and days: The optimal send time varies dramatically by audience. Test morning versus afternoon, weekday versus weekend, and specific time slots. CoSchedule’s meta-analysis found that Tuesday, Thursday, and Wednesday tend to produce the highest engagement, but your audience may be different.
- From name and email address: Test whether emails from a personal name (e.g., “Anna from MailCraft”) outperform those from a brand name (e.g., “MailCraft”). Pinpointe research shows that using a real person’s name can increase open rates by up to 35% in B2B contexts.
Start with the elements that have the highest potential impact on your primary metric, and build a testing roadmap that systematically works through each element over time.
4. Segment and size your test audience correctly
The reliability of your test results depends entirely on how you construct your test groups:
- Random assignment: Subscribers must be randomly assigned to either the A or B group. Any systematic bias (such as assigning more recent subscribers to one group) will invalidate your results. MailCraft handles this randomisation automatically.
- Equal group sizes: Both groups should be the same size to ensure a fair comparison. If your list has 10,000 subscribers, you might send Version A to 2,000 and Version B to 2,000, then deploy the winner to the remaining 6,000.
- Sufficient sample size: This is where many marketers go wrong. To detect a meaningful difference with 95% confidence, you need a minimum sample size that depends on your baseline metric and the size of the improvement you want to detect. As a rough guideline:
For a list with a baseline open rate of 20%, detecting a 2-percentage-point improvement (to 22%) with 95% confidence requires approximately 3,800 subscribers per variant. Use an online sample size calculator to determine the exact number for your situation.
- Segment-specific testing: If your audience is diverse, consider running separate tests for different segments. The winning variant for your enterprise customers may be different from the winner for small business subscribers. MailCraft lets you run segment-specific A/B tests easily.
Proper audience sizing and randomisation are the technical foundation that makes everything else in A/B testing valid. Skip this step, and your results may be misleading.
5. Create your test variants with precision
The quality of your test variants directly determines the quality of your insights:
- Version A (Control): This is your current, established approach - the email as you would normally send it. It serves as the baseline against which the new idea is measured.
- Version B (Variation): This is the challenger - the new element you believe might outperform the control. Make the change meaningful enough to potentially produce a measurable difference. Testing “Buy Now” versus “Buy Now!” (just adding an exclamation mark) is unlikely to produce statistically significant results.
- Maintain strict consistency: Everything except the single variable being tested must be identical across both versions. Same send time, same from address, same audience demographic profile, same email layout (unless layout is the variable being tested).
- Quality assurance: Before launching, have at least one other person review both versions. Check for typos, broken links, rendering issues across email clients, and any unintended differences. A broken link in one version could completely invalidate your test results.
Attention to detail in variant creation is what separates rigorous, trustworthy tests from experiments that produce misleading conclusions.
6. Determine the right sample size and test duration
Knowing when to end a test - and when to keep it running - is one of the most important skills in A/B testing:
- Use sample size calculators: Tools like Optimizely’s sample size calculator or Evan Miller’s calculator let you input your baseline conversion rate, the minimum detectable effect (the smallest improvement you care about), and your desired confidence level to get the exact sample size you need.
- Set a minimum test duration: Even if you reach your required sample size quickly, run the test for at least one full business cycle (typically 7 days) to account for day-of-week variations in subscriber behaviour.
- Account for external factors: Avoid running tests during holidays, major sales events, or any period where subscriber behaviour may be atypical. These anomalies can skew your results. If you launch a test on Black Friday, the results will not reflect normal behaviour.
- Resist the temptation to peek: Checking results hourly and stopping the test as soon as one variant “looks like it’s winning” is the fastest way to generate false positives. Commit to your predetermined sample size and duration before analysing results.
Discipline in test timing is what separates actionable insights from statistical noise.
7. Leverage MailCraft’s built-in A/B testing capabilities
You do not need to manage A/B tests manually - MailCraft’s testing features handle the heavy lifting:
- Easy test setup: Create A/B tests directly within MailCraft’s campaign builder. Select the element you want to test, create your variants, and define your success metric - all in a few clicks.
- Automated random distribution: MailCraft automatically and randomly assigns subscribers to each variant group, ensuring unbiased results without any manual list manipulation.
- Real-time performance tracking: Monitor how each variant is performing as results come in, with live dashboards showing open rates, click rates, and conversion rates for both versions.
- Automatic winner deployment: Configure MailCraft to automatically send the winning variant to the remainder of your list once statistical significance is reached. This feature maximises the impact of every test by ensuring the best-performing version reaches the largest possible audience.
- Test history and documentation: MailCraft stores the results of all previous tests, building a knowledge base that informs future testing decisions. Over time, this institutional memory becomes one of your most valuable marketing assets.
By leveraging these features, you can run rigorous A/B tests without needing a data scientist on your team.
8. Analyse results rigorously - not just the headline numbers
When your test concludes, thorough analysis is what turns raw data into actionable intelligence:
- Confirm statistical significance: Before drawing any conclusions, verify that the difference between variants is statistically significant (typically at the 95% confidence level). A result showing Variant B has a 22% open rate versus Variant A’s 21% means nothing if the sample size is too small for that difference to be significant.
- Examine multiple metrics: While your primary metric should drive the decision, review secondary metrics for a complete picture. A subject line that dramatically increases open rates but leads to higher unsubscribe rates may not be the right choice overall.
- Segment the results: Look at how different subscriber segments responded to each variant. The overall winner might not be the winner for every segment - and this insight can inform more sophisticated, segment-specific strategies.
- Calculate the business impact: Translate the percentage improvement into real numbers. “A 3% increase in click-through rate” becomes much more compelling when you calculate that it represents an additional 450 clicks per campaign, leading to an estimated 45 additional conversions worth $4,500 in revenue.
- Document everything: Record the test hypothesis, variants, sample size, duration, results, statistical significance, and your interpretation. This documentation becomes invaluable as you build a library of tested learnings specific to your audience.
Econsultancy reports that only 28% of marketers are satisfied with their conversion rates. Rigorous A/B testing is the most reliable path from dissatisfaction to improvement - but only if the analysis is done properly.
9. Implement winning variants and measure sustained impact
Identifying a winner is only valuable if you act on the finding:
- Deploy the winning variant: Update your email templates, automation workflows, and campaign defaults to incorporate the winning element. If your test showed that conversational subject lines outperform formal ones, update your subject line guidelines across all campaigns.
- Ensure brand alignment: Verify that the winning variant is consistent with your overall brand voice, visual identity, and messaging strategy. A clickbait-style subject line might win an open rate test but damage brand trust over time.
- Monitor long-term performance: A single test result is a snapshot. Continue tracking the implemented change over multiple campaigns to confirm that the improvement persists. Subscriber behaviour can shift over time due to seasonal changes, list composition changes, or broader market trends.
- Watch for novelty effects: Sometimes a new approach performs well initially simply because it is different - not because it is genuinely better. If you see performance decline after the first few deployments of a winning variant, the initial test may have captured a novelty effect rather than a sustainable improvement.
The goal is not just to win individual tests but to build a continuously improving email marketing programme where every test contributes to cumulative knowledge and better performance.
10. Build an iterative, ongoing testing programme
A/B testing is not a one-time project - it is a permanent operating discipline that drives continuous optimisation:
- Maintain a testing roadmap: Keep a prioritised backlog of test ideas, ranked by potential impact and ease of implementation. After completing one test, immediately queue the next one. High-performing email programmes typically run 2-4 A/B tests per month.
- Test systematically, not randomly: Work through email elements in a logical sequence. Start with subject lines (highest impact on opens), then move to CTAs (highest impact on clicks), then content structure, then send times. Each test builds on the learnings of previous tests.
- Re-test periodically: What works today may not work in six months. Subscriber preferences evolve, inbox algorithms change, and competitive dynamics shift. Re-test your key assumptions at least annually to ensure your strategies remain effective.
- Share learnings across the organisation: Email testing insights often apply to other channels. A subject line test that reveals your audience prefers questions over statements might also inform your social media headlines, ad copy, and landing page headlines.
HubSpot data shows that companies with a structured testing programme achieve conversion rates 2-3x higher than those that test sporadically or not at all. The compounding effect of continuous, systematic optimisation is powerful.
11. Graduate to multivariate testing for deeper insights
Once you have mastered A/B testing and built a solid foundation of learnings, you may be ready to explore multivariate testing (MVT) - a more advanced technique that tests multiple variables simultaneously:
- How MVT works: Instead of testing one variable with two variants, MVT tests combinations of multiple variables. For example, you might test two subject lines crossed with two CTA designs, creating four distinct variants (A1, A2, B1, B2). This allows you to understand not just individual effects but also interaction effects - how variables influence each other.
- When to use MVT: Multivariate testing requires significantly larger audiences to achieve statistical significance (because you are splitting traffic across more variants). It is most appropriate when you have a large list (50,000+ subscribers) and want to understand complex relationships between email elements.
- Practical applications: MVT is particularly useful for optimising landing pages linked from emails, where the combination of headline, hero image, and CTA text can produce non-obvious winning combinations that you would never discover through sequential A/B tests alone.
Think of multivariate testing as the advanced course - pursue it after you have built a strong foundation with standard A/B testing.
12. Maintain GDPR compliance throughout your testing programme
A/B testing involves processing subscriber data, which means compliance with the General Data Protection Regulation (GDPR) and other data protection laws is non-negotiable:
- Lawful basis for processing: Ensure you have a valid lawful basis (typically consent or legitimate interest) for sending marketing emails and conducting tests. A/B testing is generally considered a normal part of marketing operations and does not require separate consent, but your privacy policy should mention it.
- Transparency: Your privacy policy should inform subscribers that you may conduct testing to improve the relevance and quality of communications they receive.
- Data minimisation: Only collect and use the data you actually need for segmentation and testing. Do not accumulate unnecessary personal data.
- Anonymised reporting: When sharing test results internally, use aggregated data rather than individual subscriber-level data wherever possible.
MailCraft includes built-in GDPR compliance tools - including consent management, data export, and deletion capabilities - that help you maintain compliance while running a sophisticated testing programme.
13. Engage your team to multiply testing effectiveness
The best testing programmes are collaborative, not siloed in a single person’s inbox:
- Brainstorm test ideas across departments: Your sales team knows what objections customers raise. Your support team knows what questions come up most often. Your product team knows what features drive the most engagement. All of these insights can generate high-value test hypotheses.
- Establish a regular testing review cadence: Hold a monthly meeting to review completed tests, discuss findings, and prioritise the next round of experiments. This keeps the testing programme visible and valued across the organisation.
- Create a shared knowledge base: Maintain a centralised document or wiki that catalogues every test you have run, the results, and the business implications. Over time, this becomes a powerful institutional resource.
- Celebrate both wins and failures: A test where the variation loses is still valuable - it tells you what does not work and prevents you from investing in the wrong direction. Teams that celebrate learning (not just winning) run more tests and learn faster.
14. Extract maximum value from unsuccessful tests
Not every test will produce a winner, and that is perfectly normal. What matters is how you respond to inconclusive or negative results:
- Analyse why the variation underperformed: Was the change too subtle to produce a measurable difference? Was the hypothesis flawed? Did an external factor (like a holiday or a technical issue) contaminate the results?
- Refine your hypothesis: Use the insights from unsuccessful tests to formulate better hypotheses for the next round. If personalised subject lines did not increase opens, perhaps the personalisation was too shallow (just a first name) - test deeper personalisation next (referencing a recent purchase or browsing behaviour).
- Check for segment-level insights: An overall null result might mask strong segment-specific effects. The variation might have significantly outperformed for one segment while underperforming for another, cancelling out in the aggregate.
- Know when to accept the null result: If a well-designed, properly sized test shows no significant difference between variants, the honest conclusion is that the variable you tested does not meaningfully impact performance for your audience. Accept this and move your testing resources to higher-impact elements.
Thomas Edison famously said, “I have not failed. I have just found 10,000 ways that won’t work.” In A/B testing, every failed experiment narrows the search space and brings you closer to strategies that do work.
15. Apply email testing insights across all your marketing channels
One of the most underappreciated benefits of email A/B testing is that the insights often transfer to other channels:
- Email to social media: Subject line tests reveal what language, tone, and framing resonates with your audience. Apply winning subject line formulas to social media headlines, ad copy, and post captions.
- Email to website: CTA tests in emails can inform landing page button design. Content length tests can guide blog post formatting. Imagery tests can shape your website’s visual direction.
- Email to paid advertising: Offers and promotions that perform well in email are strong candidates for paid advertising campaigns. Conversely, ad copy that drives high click-through rates can be adapted for email subject lines.
- Build a unified customer understanding: Every test teaches you something about your audience’s preferences, decision-making processes, and communication style. These insights are not channel-specific - they reflect who your customers are. Use them everywhere.
By treating your email testing programme as a learning engine for your entire marketing operation, you multiply the return on every test you run.
Conclusion: make A/B testing your competitive advantage
A/B testing is the most reliable method available for improving email campaign performance. By following a structured process - understanding the fundamentals, defining clear objectives, choosing the right elements to test, sizing your audience correctly, creating precise variants, allowing sufficient test duration, leveraging platform tools, analysing results rigorously, implementing winners, building an iterative programme, exploring multivariate testing, maintaining compliance, engaging your team, learning from failures, and applying insights across channels - you create a compounding cycle of improvement that sets your email marketing apart from the competition.
MailCraft is designed to make A/B testing accessible and rigorous, offering an intuitive testing interface, automated random distribution, real-time analytics, automatic winner deployment, and comprehensive result documentation. Whether you are running your first subject line test or building a sophisticated multivariate testing programme, MailCraft provides the tools and infrastructure you need to turn data into decisions and decisions into results.
Start testing today. The businesses that commit to systematic, ongoing A/B testing will consistently outperform those that rely on intuition - and the gap will only widen over time.


