[{"data":1,"prerenderedAt":861},["ShallowReactive",2],{"legal-pages":3,"blog-index":13},[4,7,9,11],{"path":5,"draft":6},"\u002Flegal\u002Facceptable-use",false,{"path":8,"draft":6},"\u002Flegal\u002Fprivacy",{"path":10,"draft":6},"\u002Flegal\u002Fsubprocessors",{"path":12,"draft":6},"\u002Flegal\u002Fterms",[14,346,620],{"id":15,"title":16,"author":17,"body":18,"date":330,"description":331,"draft":6,"extension":332,"meta":333,"navigation":334,"path":335,"seo":336,"stem":337,"tags":338,"tool":343,"updated":344,"__hash__":345},"blog\u002Fblog\u002Fgmail-yahoo-bulk-sender-rules.md","Gmail and Yahoo's bulk sender rules, in plain English","Carrier Crow",{"type":19,"value":20,"toc":318},"minimark",[21,25,28,33,36,39,42,145,149,160,175,178,182,192,202,215,230,234,237,240,246,249,252,256,259,262,265,269,276,282,286,289,298,301,305,308,312],[22,23,24],"p",{},"In February 2024, Gmail and Yahoo stopped treating good sending practice as optional. Authentication, an easy way out, a ceiling on complaints: most of it had been advice for years. What changed is enforcement. It was phased in rather than switched on overnight, but the phase-in is long over, and a newsletter that misses the bar now risks being deferred, filtered or rejected.",[22,26,27],{},"Here is what the rules ask for, what each one means, and how to check yours.",[29,30,32],"h2",{"id":31},"who-counts-as-a-bulk-sender","Who counts as a bulk sender",[22,34,35],{},"Google defines a bulk sender as anyone sending close to 5,000 messages or more to personal Gmail accounts in a 24-hour period, counting everything sent from the same primary domain. Google has also said that once you cross that line, you're treated as a bulk sender from then on. One big issue is enough.",[22,37,38],{},"Yahoo announced matching requirements at the same time, and Microsoft has since set out similar authentication rules for high-volume senders to Outlook.com, Hotmail and Live addresses. If you're anywhere near 5,000 a day to Gmail, meet the bulk requirements everywhere.",[22,40,41],{},"Some rules apply to every sender, however small; others only to bulk senders:",[43,44,45,61],"table",{},[46,47,48],"thead",{},[49,50,51,55,58],"tr",{},[52,53,54],"th",{},"Requirement",[52,56,57],{},"Who it applies to",[52,59,60],{},"How to check",[62,63,64,76,92,102,113,124,134],"tbody",{},[49,65,66,70,73],{},[67,68,69],"td",{},"SPF and DKIM",[67,71,72],{},"Everyone needs one; bulk senders need both",[67,74,75],{},"DNS, and a received message",[49,77,78,86,89],{},[67,79,80,81,85],{},"A DMARC record, ",[82,83,84],"code",{},"p=none"," or stricter",[67,87,88],{},"Bulk senders",[67,90,91],{},"DNS",[49,93,94,97,99],{},[67,95,96],{},"From domain aligned with SPF or DKIM",[67,98,88],{},[67,100,101],{},"The DMARC result on a received message",[49,103,104,107,110],{},[67,105,106],{},"One-click unsubscribe, honoured within two days",[67,108,109],{},"Bulk senders, for marketing and subscribed mail",[67,111,112],{},"Message headers",[49,114,115,118,121],{},[67,116,117],{},"Spam rate below 0.3%",[67,119,120],{},"Everyone",[67,122,123],{},"Google Postmaster Tools",[49,125,126,129,131],{},[67,127,128],{},"Forward and reverse DNS for sending IPs",[67,130,120],{},[67,132,133],{},"A reverse lookup on the IP",[49,135,136,139,142],{},[67,137,138],{},"TLS",[67,140,141],{},"Everyone (Google)",[67,143,144],{},"Message details in Gmail",[29,146,148],{"id":147},"spf-and-dkim-proving-the-mail-is-yours","SPF and DKIM: proving the mail is yours",[22,150,151,155,156,159],{},[152,153,154],"strong",{},"SPF"," is a DNS record on your domain listing the servers allowed to send mail for it. A receiving server takes the domain from the envelope sender (the hidden return address that bounces go to, not the From line readers see) and checks whether the connecting server is on the list. Two classic mistakes: publishing more than one SPF record, which counts as an error rather than a merge, and chaining so many ",[82,157,158],{},"include:"," entries that you pass SPF's limit of ten DNS lookups. Either makes SPF fail while the record looks fine at a glance.",[22,161,162,165,166,170,171,174],{},[152,163,164],{},"DKIM"," is a cryptographic signature the sending server adds to each message. The matching public key lives in DNS under a ",[167,168,169],"em",{},"selector"," you or your provider choose, at an address like ",[82,172,173],{},"s1._domainkey.example.com",". Because the signature travels inside the message, DKIM survives forwarding in a way SPF doesn't.",[22,176,177],{},"Miss both and Gmail has little reason to believe the mail is from you. For a bulk sender, having only one is itself out of compliance.",[29,179,181],{"id":180},"dmarc-and-the-alignment-trap","DMARC, and the alignment trap",[22,183,184,187,188,191],{},[152,185,186],{},"DMARC"," is a TXT record at ",[82,189,190],{},"_dmarc.example.com"," telling receivers what to do with mail that claims to be from your domain but fails authentication, and where to send reports. The minimum the rules ask for is a monitoring-only policy:",[193,194,199],"pre",{"className":195,"code":197,"language":198},[196],"language-text","v=DMARC1; p=none; rua=mailto:dmarc-reports@example.com\n","text",[82,200,197],{"__ignoreMap":201},"",[22,203,204,206,207,210,211,214],{},[82,205,84],{}," means \"don't act on failures, just tell me about them\". It's a starting point: the aggregate reports it produces show every service sending as your domain, which is how you find out when it's safe to move to ",[82,208,209],{},"quarantine"," or ",[82,212,213],{},"reject",".",[22,216,217,218,221,222,225,226,229],{},"The part that catches people out is ",[152,219,220],{},"alignment",". DMARC passes only if the domain in your visible From address matches the domain that passed SPF, or the domain that signed with DKIM. By default a match at the organisational level is enough, so ",[82,223,224],{},"news.example.com"," aligns with ",[82,227,228],{},"example.com",". The trap: unless you've set up domain authentication, many providers sign with their own DKIM domain and use their own bounce domain. SPF passes, DKIM passes, and DMARC still fails, because neither passing domain is yours. Your provider's domain authentication setup fixes it by making it sign and bounce as you.",[29,231,233],{"id":232},"one-click-unsubscribe","One-click unsubscribe",[22,235,236],{},"Bulk senders must let people unsubscribe from marketing and subscribed mail (newsletters included) in one click, and must honour the request within two days. Google also wants a clearly visible unsubscribe link in the body of the message; the header doesn't replace it.",[22,238,239],{},"\"One click\" has a specific technical meaning, set out in RFC 8058. Each message carries two headers:",[193,241,244],{"className":242,"code":243,"language":198},[196],"List-Unsubscribe: \u003Chttps:\u002F\u002Fexample.com\u002Funsubscribe\u002Fabc123>\nList-Unsubscribe-Post: List-Unsubscribe=One-Click\n",[82,245,243],{"__ignoreMap":201},[22,247,248],{},"When a reader uses the unsubscribe option Gmail or Yahoo shows beside your name, the mailbox provider sends an HTTPS POST to that address and your system unsubscribes them. No login, no confirmation page. It's a POST rather than a simple visit because security software follows links on its own, and nobody should be unsubscribed by a scanner. The standard also requires both headers to be covered by the message's DKIM signature.",[22,250,251],{},"Miss this and, compliance aside, readers who can't find the way out use the one button that's always there: \"Report spam\". Which brings us to the number that matters most.",[29,253,255],{"id":254},"keep-spam-complaints-below-03","Keep spam complaints below 0.3%",[22,257,258],{},"Google wants the spam rate shown in its Postmaster Tools kept below 0.3%, and recommends staying under 0.1%. Yahoo uses the same 0.3% line. Google measures the rate against mail that reached the inbox, so 0.3% is three complaints for every thousand messages delivered there. On 20,000 Gmail inbox deliveries, that's 60 people.",[22,260,261],{},"Over the line, more of your mail goes to spam, and Google has said senders above it lose eligibility for its mitigation help until the rate comes back down. Postmaster Tools is free: you verify your domain with a DNS record, though it only shows data once you send enough mail to Gmail each day. Yahoo offers a complaint feedback loop through its Sender Hub.",[22,263,264],{},"The levers are unglamorous: confirm signups, make leaving easy, stop mailing people who never open.",[29,266,268],{"id":267},"the-plumbing-reverse-dns-and-tls","The plumbing: reverse DNS and TLS",[22,270,271,272,275],{},"Every IP address you send from needs ",[152,273,274],{},"forward and reverse DNS",": a PTR record mapping the IP to a hostname, and that hostname resolving back to the same IP. Send through SendGrid, Mailgun, Postmark or Amazon SES on their shared IPs and the provider handles this. Run your own mail server and it's on you.",[22,277,278,279,281],{},"Google also requires mail to arrive over ",[152,280,138],{},". Mainstream providers do this by default. In Gmail, a message's details read \"Standard encryption (TLS)\"; mail that arrived without it gets a red open padlock.",[29,283,285],{"id":284},"how-to-check-where-you-stand","How to check where you stand",[22,287,288],{},"Two checks cover most of it.",[22,290,291,292,297],{},"First, the DNS: SPF, DMARC, your DKIM key at its selector, and MX. MX isn't on Google's list, but replies and bounces need somewhere to go, and some receiving servers turn away mail whose return-address domain can't receive any. The ",[293,294,296],"a",{"href":295},"\u002Ftools\u002Fsender-check","sender check"," looks up all four for a domain in one go.",[22,299,300],{},"Second, a real message. Send an issue to a Gmail address, open it and choose \"Show original\". The summary at the top shows SPF, DKIM and DMARC each as PASS or FAIL. DNS tells you the records exist; only a received message tells you they line up. If DMARC says PASS, alignment is working, and the raw headers underneath show whether both unsubscribe headers made it through.",[29,302,304],{"id":303},"where-carrier-crow-fits","Where Carrier Crow fits",[22,306,307],{},"Carrier Crow isn't an email service provider. It connects to the SMTP provider you already use (SendGrid, Mailgun, Postmark, Amazon SES or any SMTP relay), so SPF, DKIM, DMARC and reverse DNS stay with your domain and your provider. The part that lives inside the message, it handles: every campaign carries the one-click List-Unsubscribe headers, whichever provider sends it.",[29,309,311],{"id":310},"try-it","Try it",[22,313,314,315,317],{},"Put your sending domain into the ",[293,316,296],{"href":295},". It runs the SPF, DKIM, DMARC and MX lookups for you, which beats reading TXT records by hand before the next issue goes out.",{"title":201,"searchDepth":319,"depth":319,"links":320},2,[321,322,323,324,325,326,327,328,329],{"id":31,"depth":319,"text":32},{"id":147,"depth":319,"text":148},{"id":180,"depth":319,"text":181},{"id":232,"depth":319,"text":233},{"id":254,"depth":319,"text":255},{"id":267,"depth":319,"text":268},{"id":284,"depth":319,"text":285},{"id":303,"depth":319,"text":304},{"id":310,"depth":319,"text":311},"2026-10-07","What Gmail and Yahoo have required of bulk senders since February 2024, what breaks if you miss a rule, and how to check your own setup.","md",{},true,"\u002Fblog\u002Fgmail-yahoo-bulk-sender-rules",{"title":16,"description":331},"blog\u002Fgmail-yahoo-bulk-sender-rules",[339,340,341,342],"deliverability","authentication","gmail","yahoo","sender-check",null,"xY1l3JQTmylDRYi-KNi2hbOChyPuNbZDKFtakNJ4xqU",{"id":347,"title":348,"author":17,"body":349,"date":330,"description":609,"draft":6,"extension":332,"meta":610,"navigation":334,"path":611,"seo":612,"stem":613,"tags":614,"tool":618,"updated":344,"__hash__":619},"blog\u002Fblog\u002Fhow-long-do-subscribers-stay.md","How long do your subscribers really stay?",{"type":19,"value":350,"toc":598},[351,354,357,361,364,367,374,378,381,384,388,391,394,465,468,475,479,482,504,508,515,521,525,528,531,535,538,551,554,580,583,587,590,592],[22,352,353],{},"It sounds like the easiest question in the newsletter business. Export everyone who has unsubscribed, subtract each signup date from the unsubscribe date, take the average. Done: your subscribers stay for, say, four months.",[22,355,356],{},"That number is wrong, in a predictable direction: too short. Subscriber lifetime feeds into what a reader is worth and what you can promise sponsors, so it's worth getting right.",[29,358,360],{"id":359},"why-the-obvious-average-comes-out-too-short","Why the obvious average comes out too short",[22,362,363],{},"The problem is who's in the calculation. Averaging the people who left ignores everyone who hasn't, which is often most of your list and, by definition, your most loyal readers. Someone who signed up three years ago and still reads every issue counts for nothing.",[22,365,366],{},"The leavers themselves are a skewed sample, too. Nobody in the export can have stayed longer than your list has existed, and someone who signed up two months ago can only appear in it if they left within two months. The calculation can see short lifetimes clearly; the long ones are mostly still in progress.",[22,368,369,370,373],{},"Statisticians call the people still subscribed ",[167,371,372],{},"censored",": you know they've stayed at least this long, but not how long they'll stay in the end. Throwing them away biases the answer. Treating them as if they'd left today is less bad, but still too short, because their clocks are still running.",[29,375,377],{"id":376},"a-worked-example","A worked example",[22,379,380],{},"Take ten subscribers. Five have left, after 1, 2, 2, 4 and 7 months. Five are still subscribed, and have been for 3, 6, 10, 14 and 20 months.",[22,382,383],{},"The leavers-only average is 3.2 months. Now let's give the five people who stayed a say.",[29,385,387],{"id":386},"what-kaplanmeier-does","What Kaplan–Meier does",[22,389,390],{},"The Kaplan–Meier estimator is the standard method for this kind of data. It walks forward through time, and at each point where someone leaves it asks: of the people still subscribed and still being observed at that moment, what share stayed? Then it multiplies those shares together.",[22,392,393],{},"The crucial rule is how it treats someone who's still subscribed. They count, in full, for as long as they've been with you. After that they quietly drop out of the group being observed (the \"at-risk\" group) without being counted as a loss.",[43,395,396,412],{},[46,397,398],{},[49,399,400,403,406,409],{},[52,401,402],{},"Month",[52,404,405],{},"Subscribed going in",[52,407,408],{},"Left that month",[52,410,411],{},"Share still subscribed",[62,413,414,427,440,453],{},[49,415,416,419,422,424],{},[67,417,418],{},"1",[67,420,421],{},"10",[67,423,418],{},[67,425,426],{},"90%",[49,428,429,432,435,437],{},[67,430,431],{},"2",[67,433,434],{},"9",[67,436,431],{},[67,438,439],{},"70%",[49,441,442,445,448,450],{},[67,443,444],{},"4",[67,446,447],{},"6",[67,449,418],{},[67,451,452],{},"58%",[49,454,455,458,460,462],{},[67,456,457],{},"7",[67,459,444],{},[67,461,418],{},[67,463,464],{},"44%",[22,466,467],{},"Look at month 4. Three people have left by then, so you might expect seven in the group, but there are six: the subscriber who joined three months ago has no history beyond month 3, so they stop counting there, without being marked as lost. Before month 7, the six-month subscriber drops out the same way. Each step multiplies the last: 90% × 7\u002F9 = 70%, then × 5\u002F6 ≈ 58%, then × 3\u002F4 ≈ 44%. Nobody leaves after month 7, so the curve stays at 44% out to month 20.",[22,469,470,471,214],{},"So the same ten people tell a different story. Instead of \"subscribers last 3.2 months\", it's \"half are gone by month 7, and 44% are still here after more than a year\". That's a different business. Ten people you can do by hand; for ten thousand, there's the ",[293,472,474],{"href":473},"\u002Ftools\u002Fsubscriber-lifetime","subscriber lifetime tool",[29,476,478],{"id":477},"reading-a-survival-curve","Reading a survival curve",[22,480,481],{},"Plot the share still subscribed against time and you get a staircase. It starts at 100% and steps down each time someone leaves. Three things to look at:",[483,484,485,492,498],"ul",{},[486,487,488,491],"li",{},[152,489,490],{},"Where it drops fastest."," A steep fall in the first few weeks suggests looking at signup and onboarding; a steady slope later on points more towards the content itself.",[486,493,494,497],{},[152,495,496],{},"Where it flattens."," A flat stretch means readers who reach that point tend to stay.",[486,499,500,503],{},[152,501,502],{},"How much data sits under the tail."," The far right of the curve rests on your oldest subscribers, often only a handful of people, so one departure moves it a long way. Treat the tail as a sketch.",[29,505,507],{"id":506},"median-lifetime-or-retention-at-fixed-points","Median lifetime, or retention at fixed points",[22,509,510,511,514],{},"The ",[152,512,513],{},"median lifetime"," is the point where the curve first falls to 50% or below: the time by which half of subscribers have left. In the example it's 7 months, more than twice the naive average. It's one intuitive number, and it isn't dragged around by a few very long or very short lifetimes the way an average is. For a list that keeps its readers it may not exist yet: if more than half are still subscribed at the longest time you can observe, the honest answer is \"longer than that\".",[22,516,517,520],{},[152,518,519],{},"Retention at fixed points"," is often more useful: the share still subscribed at 90, 180 and 365 days. It's defined as long as some of your subscribers have been around that long, it compares cleanly across months and sources, and it answers practical questions. Of the readers a promotion brings in, how many will still be there for next quarter's sponsors?",[29,522,524],{"id":523},"comparing-cohorts","Comparing cohorts",[22,526,527],{},"The method works on any group you can label. Split by signup source (your own site, a cross-promotion, a giveaway, a paid campaign) and draw a curve for each. If one source's curve falls away much faster, its headline signup numbers overstate what you're getting from it.",[22,529,530],{},"Two cautions. Compare groups only at time points both have data for: a cohort that joined in spring can't tell you its 365-day retention yet. And small groups give jumpy curves, so be wary of reading much into a gap between two cohorts of a few dozen people. For a formal comparison, the log-rank test is the standard one.",[29,532,534],{"id":533},"what-to-export","What to export",[22,536,537],{},"You need two columns per subscriber:",[483,539,540,545],{},[486,541,542],{},[152,543,544],{},"signup date",[486,546,547,550],{},[152,548,549],{},"unsubscribe date",", left blank if they're still subscribed",[22,552,553],{},"Add a column such as signup source to compare cohorts. Then check a few things:",[483,555,556,562,568,574],{},[486,557,558,561],{},[152,559,560],{},"Unsubscribed records."," Some platforms delete them or keep them elsewhere. An export of current subscribers contains no losses at all, and any method will be wildly optimistic.",[486,563,564,567],{},[152,565,566],{},"Migrated lists."," If you've changed platforms, signup dates may all be the day of the import, which resets everyone's clock. Use the original dates if you have them.",[486,569,570,573],{},[152,571,572],{},"Bounces and cleaned addresses."," Decide whether these count as leaving. They're gone, but not by choice; either answer is defensible as long as you're consistent.",[486,575,576,579],{},[152,577,578],{},"Resubscribers."," Pick a rule (first signup, or each spell as its own row) and stick to it.",[22,581,582],{},"One more thing: this measures subscription, not attention. Someone who hasn't opened anything in a year still counts as subscribed. If what you care about is engaged lifetime, define \"leaving\" as going inactive, and the same method applies.",[29,584,586],{"id":585},"in-carrier-crow","In Carrier Crow",[22,588,589],{},"Carrier Crow has a built-in Subscriber Lifetime report that uses Kaplan–Meier to give you the median lifetime and the share still subscribed at 90, 180 and 365 days.",[29,591,311],{"id":310},[22,593,594,595,597],{},"Export signup and unsubscribe dates from whichever platform you use and drop the file into the ",[293,596,474],{"href":473},". It runs entirely in your browser, so the file is never uploaded, and you'll see how long your readers really stay, counting the ones who haven't left.",{"title":201,"searchDepth":319,"depth":319,"links":599},[600,601,602,603,604,605,606,607,608],{"id":359,"depth":319,"text":360},{"id":376,"depth":319,"text":377},{"id":386,"depth":319,"text":387},{"id":477,"depth":319,"text":478},{"id":506,"depth":319,"text":507},{"id":523,"depth":319,"text":524},{"id":533,"depth":319,"text":534},{"id":585,"depth":319,"text":586},{"id":310,"depth":319,"text":311},"Averaging how long departed subscribers lasted makes your list look fickle. Survival analysis counts everyone, including the people still here.",{},"\u002Fblog\u002Fhow-long-do-subscribers-stay",{"title":348,"description":609},"blog\u002Fhow-long-do-subscribers-stay",[615,616,617],"retention","analytics","survival-analysis","subscriber-lifetime","0W_5o9C8dmetaadwnjl_uvrK7DVoEGlvfkaEQRdeVG0",{"id":621,"title":622,"author":17,"body":623,"date":330,"description":850,"draft":6,"extension":332,"meta":851,"navigation":334,"path":852,"seo":853,"stem":854,"tags":855,"tool":859,"updated":344,"__hash__":860},"blog\u002Fblog\u002Fis-your-ab-test-result-real.md","Is your sponsor's A\u002FB test result real?",{"type":19,"value":624,"toc":839},[625,628,631,635,638,641,645,648,675,683,687,690,708,711,715,718,726,729,733,736,739,798,801,805,808,811,814,818,821,824,828,831,833],[22,626,627],{},"A sponsor sends two versions of their ad. You run both in Tuesday's issue, each reader sees one, and by Thursday the numbers are in: creative A has a click-through rate of 4.1%, creative B 3.8%. The sponsor wants to know which won, and A looks like the answer.",[22,629,630],{},"It probably isn't. Not because A is worse, but because these numbers can't tell you either way. Here's how to see that for yourself, and what it takes to get a result you can stand behind. The same reasoning applies to subject line tests.",[29,632,634],{"id":633},"the-whole-difference-is-one-click","The whole difference is one click",[22,636,637],{},"Say each creative had 390 impressions. (For an email ad, an impression usually means an open of the issue it ran in.) A's 4.1% is 16 clicks. B's 3.8% is 15. The gap the sponsor is excited about is one person.",[22,639,640],{},"Even two identical creatives would never produce identical numbers. Clicks are a handful of people out of hundreds, and which handful clicks varies from send to send, the way 390 coin tosses don't land on exactly 195 heads. The question isn't whether the numbers differ. It's whether they differ by more than chance alone would explain.",[29,642,644],{"id":643},"the-two-proportion-z-test-without-the-algebra","The two-proportion z-test, without the algebra",[22,646,647],{},"That is exactly the question the standard test for comparing two rates asks. It goes like this:",[649,650,651,657,663,669],"ol",{},[486,652,653,656],{},[152,654,655],{},"Assume there's no real difference."," Pool both creatives: 31 clicks from 780 impressions, a combined rate of about 3.97%.",[486,658,659,662],{},[152,660,661],{},"Work out how much two samples of this size normally wobble."," With rates around 4% and 390 impressions each, the gap between two identical creatives typically swings by about 1.4 percentage points. That's the standard error.",[486,664,665,668],{},[152,666,667],{},"Compare the gap you saw with that wobble."," The observed gap is 0.26 points, about 0.18 standard errors. That ratio is the z-score.",[486,670,671,674],{},[152,672,673],{},"Turn it into a probability."," A z-score of 0.18 gives a p-value of about 0.85.",[22,676,677,678,682],{},"So if the two creatives were exactly as good as each other, you'd see a gap at least this large about 85% of the time. The result isn't evidence of anything. You don't need to do this by hand; the ",[293,679,681],{"href":680},"\u002Ftools\u002Fab-test-calculator","A\u002FB test calculator"," will.",[29,684,686],{"id":685},"what-p-005-does-and-doesnt-mean","What p \u003C 0.05 does and doesn't mean",[22,688,689],{},"The usual bar is p \u003C 0.05: if there were no real difference, a gap this large would turn up less than one time in twenty. That's a useful filter, and a widely misread one. A p-value under 0.05:",[483,691,692,698,703],{},[486,693,694,697],{},[152,695,696],{},"doesn't"," mean there's a 95% chance the winner really is better;",[486,699,700,702],{},[152,701,696],{}," tell you the difference is big, or worth anything to the sponsor, because a large enough sample makes even a trivial gap \"significant\";",[486,704,705,707],{},[152,706,696],{}," rule out luck: compare identical creatives twenty times and, on average, one comparison will clear the bar anyway.",[22,709,710],{},"What it does mean is narrower, and still valuable: the data would be surprising if nothing were going on.",[29,712,714],{"id":713},"confidence-intervals-tell-the-fuller-story","Confidence intervals tell the fuller story",[22,716,717],{},"A p-value squeezes the data into yes or no. A confidence interval shows the range of rates the data is consistent with. For our example, at 95%:",[483,719,720,723],{},[486,721,722],{},"Creative A: somewhere between about 2.5% and 6.6%",[486,724,725],{},"Creative B: somewhere between about 2.3% and 6.2%",[22,727,728],{},"And for the difference between them, anything from B ahead by 2.5 points to A ahead by 3.0. Put like that, nobody would declare a winner. The intervals are wide because the sample is small, and the honest summary is \"both creatives are somewhere around 4%\".",[29,730,732],{"id":731},"how-many-impressions-you-actually-need","How many impressions you actually need",[22,734,735],{},"The wobble shrinks with the square root of the sample size, which is the unkind part: to halve the noise, you need four times the impressions. And the finer the difference you want to detect, the faster the requirement grows. Halve the gap and you need roughly four times the sample again.",[22,737,738],{},"Here's what the standard two-proportion sample-size calculation gives at 95% confidence and 80% power (an 80% chance of detecting the difference, if it's real):",[43,740,741,754],{},[46,742,743],{},[49,744,745,748,751],{},[52,746,747],{},"One creative's CTR",[52,749,750],{},"The other's",[52,752,753],{},"Impressions needed per creative",[62,755,756,767,778,788],{},[49,757,758,761,764],{},[67,759,760],{},"3.8%",[67,762,763],{},"4.1%",[67,765,766],{},"about 66,000",[49,768,769,772,775],{},[67,770,771],{},"4.0%",[67,773,774],{},"4.4%",[67,776,777],{},"about 39,500",[49,779,780,782,785],{},[67,781,771],{},[67,783,784],{},"5.0%",[67,786,787],{},"about 6,700",[49,789,790,792,795],{},[67,791,771],{},[67,793,794],{},"6.0%",[67,796,797],{},"about 1,900",[22,799,800],{},"Turn it round: with 1,000 impressions per creative and a 4% baseline, the smallest difference you'd reliably detect is a jump to roughly 6.8%, a 70% improvement. If two creatives differ only slightly, a single issue of a mid-sized newsletter may simply not have enough readers to tell them apart. Better to know that before the test than after.",[29,802,804],{"id":803},"the-peeking-problem","The peeking problem",[22,806,807],{},"Opens and clicks trickle in over days, so it's tempting to keep refreshing the dashboard and stop the moment the result turns significant. Don't. Every look is another chance for noise to cross the line, and stopping at the first crossing means you keep the lucky readings and never see them fade.",[22,809,810],{},"The effect is large. Compare two identical creatives, check ten times as the data arrives, stop at the first p \u003C 0.05, and you'll crown a false winner about one time in five rather than one in twenty. Check twenty times and it's about one in four.",[22,812,813],{},"The fix is dull: decide the sample size up front, then judge the test once, when it's reached. There are methods designed for watching a test continuously, known as sequential tests, but they work by demanding stronger evidence at each look, for exactly this reason.",[29,815,817],{"id":816},"same-reader-same-creative","Same reader, same creative",[22,819,820],{},"One condition has nothing to do with the maths and can quietly wreck it: each reader should always see the same variant. If the variant is drawn at random every time the email renders, a reader who opens twice, or gets a resend, can see both. A click can no longer be credited cleanly to either, and the two groups stop being separate. The robust approach is to derive the variant from something stable, such as the subscriber's ID combined with the test, so the same person lands in the same group every time.",[22,822,823],{},"It's also worth making sure both sides are counting people. Security scanners that follow every link in a message add clicks no human made, and that noise lands on both creatives.",[29,825,827],{"id":826},"how-carrier-crow-handles-it","How Carrier Crow handles it",[22,829,830],{},"Carrier Crow's sponsor ad A\u002FB reporting is built around these limits. It shows no rate at all for a creative with fewer than 200 impressions, won't name a winner until every creative has at least 1,000, and names one only when a two-proportion test gives p \u003C 0.05. Which creative a reader sees is decided by a fixed rule from the reader and the ad slot, so they get the same one on the original send and on any resend. Opens and clicks are scored for bots as they arrive, and security scanners are left out of the rates by default.",[29,832,311],{"id":310},[22,834,835,836,838],{},"Before you report back to a sponsor, put each creative's impressions and clicks into the ",[293,837,681],{"href":680}," and see whether the gap clears the bar. It takes a few seconds, and it's a lot less awkward than explaining next month why the winner stopped winning.",{"title":201,"searchDepth":319,"depth":319,"links":840},[841,842,843,844,845,846,847,848,849],{"id":633,"depth":319,"text":634},{"id":643,"depth":319,"text":644},{"id":685,"depth":319,"text":686},{"id":713,"depth":319,"text":714},{"id":731,"depth":319,"text":732},{"id":803,"depth":319,"text":804},{"id":816,"depth":319,"text":817},{"id":826,"depth":319,"text":827},{"id":310,"depth":319,"text":311},"Why 4.1% against 3.8% on a few hundred impressions is noise, and how to tell a real difference between two sponsor creatives from luck.",{},"\u002Fblog\u002Fis-your-ab-test-result-real",{"title":622,"description":850},"blog\u002Fis-your-ab-test-result-real",[856,857,858],"sponsorship","ab-testing","statistics","ab-test-calculator","d-r3wGIPFl1bUs71kaGqgzvroVD-bDVrsN7eStuq4Lw",1791479469523]