Most teams know they should A/B test their cold email. Few do it, because a clean test is fiddly: split the list fairly, keep every lead on one version, wait for enough replies, then decide without fooling yourself. PitchButler’s A/B studio does that bookkeeping for you. Your job is to feed it a good question.
This guide walks through exactly what you need to create an A/B campaign, every field in the setup, how PitchButler assigns and judges your variants, and the practices that separate a useful test from a coin flip. If you want the theory first, read how to A/B test cold emails without fooling yourself.
Before you start: the checklist
Have these ready and setup takes about ten minutes:
- At least one connected Gmail mailbox with SPF, DKIM and DMARC passing. Mailboxes with failing domain authentication get no new leads, so a test on a broken domain never starts. See the SPF, DKIM and DMARC guide.
- Leads, either an imported CSV list or an Ideal Customer Profile so PitchButler can discover them. Only leads that have never been contacted and are not opted out join the test.
- One hypothesis. Write it down in a sentence: “A question subject line gets more positive replies than a statement subject line.” If you cannot write that sentence, you are not ready to test.
- Two to five versions of your first email, each with a subject line and a body, differing in the one thing your hypothesis is about.
- Your context for the AI: what you sell, what the AI may and may not say, when to hand a lead to you, and your signature.
Two ways to start a test
Open A/B studio from the Campaigns area. You have two options.
New A/B test creates a fresh campaign built for testing. Use it when you are starting outreach to a new segment.
Test new copy on a running campaign turns an existing campaign into a test. Its current first email becomes variant A, “the incumbent”, unchanged, and you add one to four challengers (B to E). Leads that were already contacted stay out of the test, so only fresh leads are compared. This is the best way to challenge copy you already rely on. A campaign can only run one test in its lifetime; to test again after a winner is promoted, duplicate the campaign.
The rest of this guide follows the New A/B test wizard, which has four steps: Basics, Lead Lists, AI Config and Review.
Step 1: Campaign basics
| Field | What to enter | Rules and defaults |
|---|---|---|
| Campaign Name | Something that states the test, e.g. “SaaS CTOs: question vs statement subject” | Required, up to 255 characters |
| Timezone | The zone your prospects live in | Defaults to your browser’s zone. Send times follow this clock |
| Leads to process per day | How many new leads the campaign starts per day | Default 50. Each mailbox’s daily limit still applies |
| Send window start / end | The hours emails may go out | Default 07:00 to 17:00. Start must be before end |
| Send days | Which weekdays to send | Default Monday to Friday. At least one day |
| Sending accounts | “Use all connected accounts” or “Only the accounts I select” | “All” also picks up mailboxes you connect later |
A few notes that matter for a test:
- Leads per day decides how fast you get an answer. More on the math below. Do not set it higher than your mailboxes can carry; the wizard shows your total capacity and warns you if you exceed it. If you are unsure, read how many cold emails to send per day.
- Leave “24/7 mode (testing)” off for real outreach. It exists for trying the system on yourself, and emails at 3 a.m. look like a machine sent them.
- Each lead keeps one mailbox for the whole conversation, and mailboxes are rotated by their daily limit, so a new, warming mailbox takes a smaller share.
Step 2: Lead lists
Tick the imported lists you want in the test, or import a CSV first. This step is optional: if you skip it, give PitchButler an Ideal Customer Profile in step 3 and it will discover leads for you.
Keep the audience one segment. If you mix Swedish agency owners with US enterprise CTOs, a winning variant may only be winning because it happened to land with the easier crowd. See how to define your ideal customer profile.
Step 3: AI configuration and your variants
This is where most of the work is. The fields in order:
- Knowledge Base: what you sell, who it is for, proof points, pricing basics. The AI uses this to write and to answer replies.
- AI Allowed Actions and AI Forbidden Actions: for example “may offer a 20 minute call” and “never promise a discount”. If an AI draft breaks a forbidden rule it is rewritten; if your own variant text breaks one, the lead is parked and you are alerted instead of sending.
- Handoff Triggers: one per line, the moments a human should take over, such as “asks for pricing” or “wants a demo”.
- Disqualification Criteria: who is not a fit, so the AI stops politely.
- Email Signature: available as the
{signature}variable. If a body does not include it, the signature is appended automatically. - Booking link: optional, must start with https://.
- Generic inboxes (info@, kontakt@ …): send normally, ask for the right contact, or skip them entirely.
- Ideal Customer Profile (optional): industries, geographies, technographics, role titles, employee and revenue ranges, and exclusions.
Your copy variants
Under Your Copy Variants you get tabs for Variant A to E. You need at least two and can have up to five. “+ Add variant” starts the new variant as a copy of the one you are on, which is exactly what you want: copy, then change one thing.
Each variant has two required fields:
- Subject Line, up to 500 characters.
- Email Body.
What is being tested is the first email of the sequence. Follow-ups are the same for everyone, which keeps the comparison clean.
One detail is worth understanding before you decide what to test. The subject line goes out exactly as you wrote it (with variables filled in). The body is personalized per lead: PitchButler’s AI uses your variant as the brief and rewrites it around that lead’s research. When there is no research for a lead, your body goes out as written. So a subject line test compares two exact subject lines, while a body test compares two messages, angles or asks, not two exact sets of words. Both are useful; just phrase your hypothesis accordingly.
Template variables
Variables work in both subject and body. Write them in single braces:
- Contact:
{first_name},{last_name},{contact_name},{title},{email} - Company:
{company_name},{website},{domain},{city} - Research:
{personalization_hooks},{suggested_angle},{pain_points},{company_summary} - Your own:
{signature}, plus any column from your CSV, e.g.{phone}
Add a fallback with |default:, for example Hi {first_name|default:there}, which becomes “Hi there,” when the first name is missing. Without a default, a missing first name falls back to the company name.
PitchButler checks every variant before it lets you save. It blocks a variant when:
- the subject or body would come out empty, or
- a placeholder is malformed, such as
{first name},{ company },{first-name}or leftover double braces.
The error names the exact token, so fix it and save again. The same check runs at send time as a safety net.
Follow-ups and the rest
- Follow-up Behavior is on by default: up to 4 follow-ups, 5 days apart, with optional instructions for tone and content. Our follow-up sequence guide covers what good spacing looks like.
- Conversation playbook (optional): up to 10 stages that guide how the AI handles a conversation. Every stage with a name needs an instruction.
- Qualify & alert only: the AI never replies on its own; it qualifies and alerts you.
Step 4: Review and launch
Review the summary, fix anything with the Edit links, and choose Create & Launch or Create Campaign to launch later. Launch needs at least one mailbox that can send and at least one lead, or an ICP to discover them.
If Approval mode is on in your Settings, every first email, from every variant, waits in your approvals inbox and exactly the text you approve is sent. That is a good way to sanity check the first handful of each variant.
How PitchButler runs the test
Assignment. Each new lead gets the active variant with the fewest leads so far. That keeps the variants evenly sized even when leads arrive at uneven times, and it is decided once: a lead never switches variant.
Pausing. Pausing a variant stops new leads from getting it; leads already on it continue. You cannot pause the last active variant.
The metric. Winners are judged on positive reply rate: replies the AI qualifier reads as interested or curious, with a booked meeting counting double. If no replies have been qualified yet, plain reply rate is used. Bounces and auto-replies never count as replies. Opens are ignored, for the reasons in cold email metrics that matter.
The minimum sample. Nothing is called before 100 sends per variant. Until then the results page says the test is still collecting, with a progress bar and an estimated finish date based on your recent pace.
The verdict. After that, PitchButler compares the leader against the runner-up with a standard two-proportion significance test. Only when the gap is unlikely to be noise (p below 0.05) does it declare a winner. Otherwise it says “No clear winner yet”, which is an honest answer, not a failure.
The results page shows, per variant: leads, sent, bounced, replied, positive, hot and won, with actions to view, edit, spam-check, pause or promote. A bounce rate of 5 percent or more (after 20 sends) is flagged in red with “check copy”. Leads emailed before the test started appear in a separate “Unassigned” row so they never muddy the comparison.
Promotion. When a variant wins you get an alert. Promoting makes the winner the campaign’s first email, archives the variants and turns the campaign back into a normal one; leads not yet emailed get the winner. By default PitchButler promotes significant winners automatically (at most once a week per campaign, and never again on a campaign where you undid one). You can switch this off with “Promote A/B winners” in Settings. A manual promotion is final.
How long will my test take?
Do the math before you launch, so you are not tempted to peek and quit early.
Each variant needs 100 first emails. With two variants that is 200 new leads; with five it is 500. At 20 leads per day on weekdays:
- 2 variants: about 10 sending days
- 3 variants: about 15 sending days
- 5 variants: about 25 sending days
Then allow a few more days for replies to arrive. This is why two or three variants beats five for most teams: every extra variant stretches the test.
Also be realistic about what 100 sends can detect. If positive reply rates are a few percent, only a large difference will clear the significance bar at this size. A test that ends in “No clear winner” has still taught you something: the change you made does not matter much, so test something bolder next.
Best practices
1. Change one thing. Subject line, opener, angle or call to action. Pick one. If B changes the subject and the ask, a win tells you nothing about which one worked.
2. Test big ideas before small words. “Question vs statement subject” or “pain angle vs proof angle” can move results. “Quick question” vs “Quick thought” almost never will, and you will not have the volume to detect it.
3. Start with the subject line. It is sent exactly as written, it gates every open, and it gives the cleanest test in the product. Our subject line guide has ideas worth testing.
4. Use body tests for message and ask. Since the body is personalized per lead, test what the email is about or asks for, not comma placement. Good body tests: a problem-first versus a result-first opening (opening lines), or a meeting ask versus an interest check (calls to action).
5. Keep length and personalization equal. If A is 60 words and B is 150, you are testing length whether you meant to or not. Use the same variables in every variant.
6. One segment per test. Same audience, same mailboxes, same send window. Everything except the variant should be identical.
7. Do not edit a variant mid-test. Edits apply from the next send, so the variant’s numbers become a blend of two emails. If you must change copy, pause the variant and read its numbers as they stood.
8. Do not stop early. The first 30 sends always look dramatic. Let the test reach 100 per variant and the verdict. The estimated finish date is there to help you be patient.
9. Watch bounces, not just replies. A variant that bounces more usually has a data or deliverability problem, not a copy advantage. Use “Check spam risk” on any variant flagged red, and see how to reduce your bounce rate.
10. Protect your sender reputation. Keep leads per day within what your mailboxes can carry, and never raise volume to finish a test faster. A burned domain costs far more than a slow test. See the sender reputation guide.
11. Compound your winners. Promote, then challenge the new champion with the next idea using “Test new copy on a running campaign” on a duplicate. Subject line this month, opener next month, ask after that. That is how a sequence gets sharper every quarter.
12. Keep a test log. One line per test: hypothesis, variants, result, decision. After five tests you will know more about your market than any template library can tell you.
Quick reference
- 2 to 5 variants (A to E), subject and body required for each
- Only the first email is tested; follow-ups are shared
- Subject sent exactly as written; body personalized per lead from your variant
- Leads are split evenly and never switch variant
- Judged on positive reply rate, meetings count double
- Minimum 100 sends per variant, winner at p below 0.05
- Auto-promotion on by default, switch in Settings
- One test per campaign; duplicate to test again
Ready to run your first one? Open A/B studio, write down your hypothesis, and let the data settle the argument.