[{"data":1,"prerenderedAt":276},["ShallowReactive",2],{"legal-pages":3,"guide:\u002Fblog\u002Fis-your-ab-test-result-real":13},[4,7,9,11],{"path":5,"draft":6},"\u002Flegal\u002Facceptable-use",false,{"path":8,"draft":6},"\u002Flegal\u002Fprivacy",{"path":10,"draft":6},"\u002Flegal\u002Fsubprocessors",{"path":12,"draft":6},"\u002Flegal\u002Fterms",{"id":14,"title":15,"author":16,"body":17,"date":261,"description":262,"draft":6,"extension":263,"meta":264,"navigation":265,"path":266,"seo":267,"stem":268,"tags":269,"tool":273,"updated":274,"__hash__":275},"blog\u002Fblog\u002Fis-your-ab-test-result-real.md","Is your sponsor's A\u002FB test result real?","Carrier Crow",{"type":18,"value":19,"toc":248},"minimark",[20,24,27,32,35,38,42,45,74,83,87,90,109,112,116,119,127,130,134,137,140,205,208,212,215,218,221,225,228,231,235,238,242],[21,22,23],"p",{},"A sponsor sends two versions of their ad. You run both in Tuesday's issue, each reader sees one, and by Thursday the numbers are in: creative A has a click-through rate of 4.1%, creative B 3.8%. The sponsor wants to know which won, and A looks like the answer.",[21,25,26],{},"It probably isn't. Not because A is worse, but because these numbers can't tell you either way. Here's how to see that for yourself, and what it takes to get a result you can stand behind. The same reasoning applies to subject line tests.",[28,29,31],"h2",{"id":30},"the-whole-difference-is-one-click","The whole difference is one click",[21,33,34],{},"Say each creative had 390 impressions. (For an email ad, an impression usually means an open of the issue it ran in.) A's 4.1% is 16 clicks. B's 3.8% is 15. The gap the sponsor is excited about is one person.",[21,36,37],{},"Even two identical creatives would never produce identical numbers. Clicks are a handful of people out of hundreds, and which handful clicks varies from send to send, the way 390 coin tosses don't land on exactly 195 heads. The question isn't whether the numbers differ. It's whether they differ by more than chance alone would explain.",[28,39,41],{"id":40},"the-two-proportion-z-test-without-the-algebra","The two-proportion z-test, without the algebra",[21,43,44],{},"That is exactly the question the standard test for comparing two rates asks. It goes like this:",[46,47,48,56,62,68],"ol",{},[49,50,51,55],"li",{},[52,53,54],"strong",{},"Assume there's no real difference."," Pool both creatives: 31 clicks from 780 impressions, a combined rate of about 3.97%.",[49,57,58,61],{},[52,59,60],{},"Work out how much two samples of this size normally wobble."," With rates around 4% and 390 impressions each, the gap between two identical creatives typically swings by about 1.4 percentage points. That's the standard error.",[49,63,64,67],{},[52,65,66],{},"Compare the gap you saw with that wobble."," The observed gap is 0.26 points, about 0.18 standard errors. That ratio is the z-score.",[49,69,70,73],{},[52,71,72],{},"Turn it into a probability."," A z-score of 0.18 gives a p-value of about 0.85.",[21,75,76,77,82],{},"So if the two creatives were exactly as good as each other, you'd see a gap at least this large about 85% of the time. The result isn't evidence of anything. You don't need to do this by hand; the ",[78,79,81],"a",{"href":80},"\u002Ftools\u002Fab-test-calculator","A\u002FB test calculator"," will.",[28,84,86],{"id":85},"what-p-005-does-and-doesnt-mean","What p \u003C 0.05 does and doesn't mean",[21,88,89],{},"The usual bar is p \u003C 0.05: if there were no real difference, a gap this large would turn up less than one time in twenty. That's a useful filter, and a widely misread one. A p-value under 0.05:",[91,92,93,99,104],"ul",{},[49,94,95,98],{},[52,96,97],{},"doesn't"," mean there's a 95% chance the winner really is better;",[49,100,101,103],{},[52,102,97],{}," tell you the difference is big, or worth anything to the sponsor, because a large enough sample makes even a trivial gap \"significant\";",[49,105,106,108],{},[52,107,97],{}," rule out luck: compare identical creatives twenty times and, on average, one comparison will clear the bar anyway.",[21,110,111],{},"What it does mean is narrower, and still valuable: the data would be surprising if nothing were going on.",[28,113,115],{"id":114},"confidence-intervals-tell-the-fuller-story","Confidence intervals tell the fuller story",[21,117,118],{},"A p-value squeezes the data into yes or no. A confidence interval shows the range of rates the data is consistent with. For our example, at 95%:",[91,120,121,124],{},[49,122,123],{},"Creative A: somewhere between about 2.5% and 6.6%",[49,125,126],{},"Creative B: somewhere between about 2.3% and 6.2%",[21,128,129],{},"And for the difference between them, anything from B ahead by 2.5 points to A ahead by 3.0. Put like that, nobody would declare a winner. The intervals are wide because the sample is small, and the honest summary is \"both creatives are somewhere around 4%\".",[28,131,133],{"id":132},"how-many-impressions-you-actually-need","How many impressions you actually need",[21,135,136],{},"The wobble shrinks with the square root of the sample size, which is the unkind part: to halve the noise, you need four times the impressions. And the finer the difference you want to detect, the faster the requirement grows. Halve the gap and you need roughly four times the sample again.",[21,138,139],{},"Here's what the standard two-proportion sample-size calculation gives at 95% confidence and 80% power (an 80% chance of detecting the difference, if it's real):",[141,142,143,159],"table",{},[144,145,146],"thead",{},[147,148,149,153,156],"tr",{},[150,151,152],"th",{},"One creative's CTR",[150,154,155],{},"The other's",[150,157,158],{},"Impressions needed per creative",[160,161,162,174,185,195],"tbody",{},[147,163,164,168,171],{},[165,166,167],"td",{},"3.8%",[165,169,170],{},"4.1%",[165,172,173],{},"about 66,000",[147,175,176,179,182],{},[165,177,178],{},"4.0%",[165,180,181],{},"4.4%",[165,183,184],{},"about 39,500",[147,186,187,189,192],{},[165,188,178],{},[165,190,191],{},"5.0%",[165,193,194],{},"about 6,700",[147,196,197,199,202],{},[165,198,178],{},[165,200,201],{},"6.0%",[165,203,204],{},"about 1,900",[21,206,207],{},"Turn it round: with 1,000 impressions per creative and a 4% baseline, the smallest difference you'd reliably detect is a jump to roughly 6.8%, a 70% improvement. If two creatives differ only slightly, a single issue of a mid-sized newsletter may simply not have enough readers to tell them apart. Better to know that before the test than after.",[28,209,211],{"id":210},"the-peeking-problem","The peeking problem",[21,213,214],{},"Opens and clicks trickle in over days, so it's tempting to keep refreshing the dashboard and stop the moment the result turns significant. Don't. Every look is another chance for noise to cross the line, and stopping at the first crossing means you keep the lucky readings and never see them fade.",[21,216,217],{},"The effect is large. Compare two identical creatives, check ten times as the data arrives, stop at the first p \u003C 0.05, and you'll crown a false winner about one time in five rather than one in twenty. Check twenty times and it's about one in four.",[21,219,220],{},"The fix is dull: decide the sample size up front, then judge the test once, when it's reached. There are methods designed for watching a test continuously, known as sequential tests, but they work by demanding stronger evidence at each look, for exactly this reason.",[28,222,224],{"id":223},"same-reader-same-creative","Same reader, same creative",[21,226,227],{},"One condition has nothing to do with the maths and can quietly wreck it: each reader should always see the same variant. If the variant is drawn at random every time the email renders, a reader who opens twice, or gets a resend, can see both. A click can no longer be credited cleanly to either, and the two groups stop being separate. The robust approach is to derive the variant from something stable, such as the subscriber's ID combined with the test, so the same person lands in the same group every time.",[21,229,230],{},"It's also worth making sure both sides are counting people. Security scanners that follow every link in a message add clicks no human made, and that noise lands on both creatives.",[28,232,234],{"id":233},"how-carrier-crow-handles-it","How Carrier Crow handles it",[21,236,237],{},"Carrier Crow's sponsor ad A\u002FB reporting is built around these limits. It shows no rate at all for a creative with fewer than 200 impressions, won't name a winner until every creative has at least 1,000, and names one only when a two-proportion test gives p \u003C 0.05. Which creative a reader sees is decided by a fixed rule from the reader and the ad slot, so they get the same one on the original send and on any resend. Opens and clicks are scored for bots as they arrive, and security scanners are left out of the rates by default.",[28,239,241],{"id":240},"try-it","Try it",[21,243,244,245,247],{},"Before you report back to a sponsor, put each creative's impressions and clicks into the ",[78,246,81],{"href":80}," and see whether the gap clears the bar. It takes a few seconds, and it's a lot less awkward than explaining next month why the winner stopped winning.",{"title":249,"searchDepth":250,"depth":250,"links":251},"",2,[252,253,254,255,256,257,258,259,260],{"id":30,"depth":250,"text":31},{"id":40,"depth":250,"text":41},{"id":85,"depth":250,"text":86},{"id":114,"depth":250,"text":115},{"id":132,"depth":250,"text":133},{"id":210,"depth":250,"text":211},{"id":223,"depth":250,"text":224},{"id":233,"depth":250,"text":234},{"id":240,"depth":250,"text":241},"2026-10-07","Why 4.1% against 3.8% on a few hundred impressions is noise, and how to tell a real difference between two sponsor creatives from luck.","md",{},true,"\u002Fblog\u002Fis-your-ab-test-result-real",{"title":15,"description":262},"blog\u002Fis-your-ab-test-result-real",[270,271,272],"sponsorship","ab-testing","statistics","ab-test-calculator",null,"d-r3wGIPFl1bUs71kaGqgzvroVD-bDVrsN7eStuq4Lw",1791479469524]