Skip to main content

Ad Testing

The SocialBuster teamPublished Sep 26, 20266 min read

What counts as enough data in a Meta ad test

The spend, conversion and impression bars a Meta ad test has to clear before a finding counts, and what a thin sample looks like just under them.

Measurewhat your ads already proved
Createthe next ad, from your own facts
Testguardrailed, and you confirm every change on Meta
Table of contents

Two bars, not one

A Meta account throws off small numbers constantly: a headline with a double-digit return on a few dollars of spend, a hook type that looks 40 percent stronger after two days of delivery. Almost none of it is real. The question is not whether a number looks good, it is whether there was enough delivery behind it to trust the number at all, and that turns out to be two different questions rather than one.

Comparing a creative attribute, a hook type against another hook type, a headline against the account average, needs at least $100 of spend and 5 conversions behind each side. That bar exists because the thing being compared is a rate built on completed purchases or leads, and a rate built on one or two conversions moves violently with the next single sale. Reading a single ad's own trend, whether this week looks different from last week, is a narrower question with a lower floor: $20 of spend and 1,000 impressions in the week. That bar is about clicks and frequency, not purchases, so it can be cleared with volume alone, well before an ad has produced enough conversions to be judged on ROAS or CPA.

Below either bar, the honest answer is not a guess dressed up as a finding. It is a plain sentence naming the actual numbers: how much was spent, how many conversions came in, and how far that sits from the bar.

The lucky purchase

Here is what sits just under both bars in practice. A test ad ran for a single day, spent $10.36 against 798 impressions, and produced exactly one conversion worth $147.22. Divide revenue by spend and the answer is a 14.21x return, a number that would sit at the top of almost any account's leaderboard.

Fourteen times the money back looks like the best ad in the account. It is one purchase away from being the worst.

That is the whole mechanic of a lucky small sample. The ratio is not wrong, $147.22 really did come back on $10.36 really spent, but the ratio is carried entirely by a single event. Had that one order landed a day later, or gone to a competitor, or simply not happened, the same ad would show a flat $0.00 in revenue and nothing else about it would be different. A number with that little weight behind it cannot be compared to anything, because comparing it means trusting a coin flip to mean something.

Fourteen times the money back looks like the best ad in the account. It is one purchase away from being the worst.

Question being askedThe barThis ad's numbersVerdict
Compare this ad's style to another$100 spend, 5 conversions$10.36 spend, 1 conversionBelow the bar
Read this ad's own weekly trend$20 spend, 1,000 impressions in 7 days$0.00 spend, 0 impressions in the last 7 daysBelow the bar

Fig. 1 · the $100/5-conversion and $20/1,000-impression bars against one real ad, Northwind Sleep demo workspace, fictional data

Clearing ahead, not just different

There is a second bar, for a different failure mode. Say two hook types both clear $100 and 5 conversions on their own, so both are individually readable. That is not the same as one of them being the winner. Real delivery is noisy enough that two adequately sized groups can sit close together purely by chance, and whichever one happens to be ahead this week can just as easily be behind next week.

To be called a winner rather than merely different, a hook type or an opening line has to beat the next best one by at least 15 percent on the metric being compared, not just edge past it. Short of that margin, the honest read is that the two are statistically too close to call, even though both individually have enough data to read. This is the same kind of judgment as the sufficiency bar, applied one level up: enough data to read a number is not the same as enough of a gap to act on it.

Below the bar is not the same as losing

An ad below the delivery floor has not been tested and lost. It has simply not been tested. Meta's own delivery system decides which ads get spend, and it can starve a perfectly reasonable ad of budget for reasons that have nothing to do with the creative, before anyone gets a fair read on it.

In the demo workspace behind this site, seven ads sit below a delivery threshold of $346.57 (the greater of $20 or 10 percent of the account's median ad spend), together accounting for just $57.00 of lifetime spend across all seven combined. That is not $57.00 of proof that seven ideas failed. It is $57.00 of evidence that the algorithm never gave any of them a real chance, evidence that only cleared roughly eight dollars a piece.

The practical rule below the bar is simple to state and easy to skip in practice: do not rank it, do not retire it, and do not average it into anything else. Give it a fair, bounded test, or wait for more delivery, and only then ask what the number means.

Where the numbers come from

The $100 spend and 5 conversion bar, and the $20 spend and 1,000 impression weekly bar, are SocialBuster's own product thresholds, the same ones Analytics, the AI Strategist and Winner Revival apply before printing a finding. The 15 percent margin is the same rule the Strategist and Lead-In Lab use when ranking hook types and opening lines against each other. These are product rules, not a general statistics claim, and they hold regardless of which account they run on.

The single test ad's numbers ($10.36 spend, 798 impressions, one $147.22 conversion, a 14.21x raw return) and the rescue figures (seven candidates, $57.00 combined lifetime spend, a $346.57 threshold) come from the Northwind Sleep demo workspace, a fictional sleep brand SocialBuster's own screens run on for this site. They show how the product reads a thin sample, not a result from any real account.

More from Proof