Foundgrove
← All posts

Paid Ads · 9 min read

Incrementality Testing: Did Your Ads Actually Cause Sales?

Summary

Attribution says which click got credit. Incrementality says whether you would have lost the revenue. How to run a geo holdout test on a real budget.

By Hyder Shah, Founder & CEO · Published July 13, 2026 · Updated July 13, 2026

Your dashboard says Google Ads produced 62 leads last month. Your agency says that's a 6x return. Neither number answers the question you actually asked, which is: if I turned this off tomorrow, would the phone stop ringing?

That question has an answer. It's called incrementality, and you can test it with two metros, a calendar, and the discipline to leave the test alone for eight weeks. No data science team. No media mix model. No new software.

What is incrementality testing, and how is it different from attribution?

Incrementality testing measures how much revenue disappears when you switch a channel off; attribution only divides credit among the clicks that were already there. Attribution is bookkeeping. Incrementality is a causal experiment with a control group — a set of customers or a market that does not see the ads, so you have something to compare against.

The gap between the two is not small. Blake, Nosko and Tadelis ran a series of large-scale field experiments at eBay and reported that returns from paid search are a fraction of conventional non-experimental estimates, and that brand-keyword ads had no measurable short-term benefit at all (NBER Working Paper 20171, published in Econometrica in 2015). Every one of those brand clicks showed up in an attribution report as a conversion. The experiment said the sales would have happened anyway.

QuestionAttribution answersIncrementality answersWhat it needs
Which touchpoint gets credit?Yes — last click, data-driven, whatever model you pickNoTracking only
Would this sale have happened without the ad?NoYesA control group
Should I cut this channel?Only by guessworkYes, with a numberA holdout test
Can it be gamed by retargeting your own buyers?EasilyNoRandomization or geo split

If you have never separated the two, start with our breakdown of why most service businesses can't tell which ad dollar made them money. Fix the tracking first. A holdout test on top of broken conversion data just gives you a confident wrong answer.

Would those customers have found you anyway?

For branded search, the honest answer is usually yes — eBay's experiments found brand-keyword ads produced no measurable short-term benefit, because the people typing your company name were already coming. That's the cannibalization problem: the ad intercepts a customer who would have clicked the organic result one line below, and then bills you for the privilege.

This does not mean brand ads are always waste. If a competitor is bidding on your name, the ad is defense, and defense has value. But it means the ROAS number on your brand campaign is the least trustworthy number in your account, and it should never be pooled with non-brand performance to justify the retainer.

The same logic applies to retargeting a warm list, and to anyone who already had your number in their phone. If a channel mostly touches people who were going to buy, its reported return is inflated by construction.

How do you run a geo holdout test with two metros?

A geo holdout is four moves: pick two comparable markets, keep the channel running in one, go dark in the other, and compare booked jobs per capita over the same window. This is the same structure Google's own research uses — Vaver and Koehler describe geo experiments as assigning non-overlapping geographic regions to a control or treatment condition and realizing the assignment through geo-targeted advertising (Measuring Ad Effectiveness Using Geo Experiments, Google, 2011).

StepWhat you doThe trap
1. Pick the pairTwo metros with similar population, similar seasonality, similar crew capacity and similar historical job volumeComparing your flagship metro to the one you opened last quarter
2. Set the baselinePull 8-12 weeks of pre-period booked jobs per metro, per week, before touching anythingUsing leads instead of booked jobs — lead quality moves between markets
3. Go darkSwitch the channel off completely in the control metro. Not 'reduce budget'. OffLeaving a Performance Max campaign quietly serving there
4. CompareBooked jobs per 100k residents, treatment vs control, test window vs baselineReading week 2 and declaring victory

Two rules make or break it. Nothing else changes during the window — no new landing page, no price change, no van wrap, no radio buy. And the control metro is genuinely dark: paused campaigns, excluded geo targets, and a check that your Local Services Ads and Performance Max are not still serving there through automatic location expansion.

If you only have one market, you can't do a geo split. What you can do is a dark period: shut the channel off for a defined block, then compare against a modeled counterfactual. Google publishes an open-source tool for exactly this — CausalImpact builds a Bayesian structural time-series model to predict how the metric would have evolved if the intervention had never happened. Its own documentation is blunt that the method needs control series that were not themselves affected by the intervention. A single-market dark period is weaker evidence than a geo split. Treat it as a strong hint, not a verdict.

How long must a holdout run before the result means anything?

Plan for 6 to 12 weeks: at least one full sales cycle after the dark switch, plus enough weeks that a normal bad month can't explain the gap. For an HVAC or roofing company where the call-to-booked-job lag is a week, 6 weeks is workable. For anything with a 60-day close — commercial contracts, B2B and SaaS pipelines — the clock doesn't even start until the pipeline has turned over once.

The other half of run length is the noise floor. If your metro does 45 jobs a month and normal month-to-month swing is 8 jobs, a test that produces a 5-job difference has told you nothing. Longer windows accumulate more jobs, which shrinks the noise relative to the effect. That's the whole reason a two-week test is theater.

Write the decision rule down before you start. “If the treatment metro books at least 15% more jobs per capita than control over 8 weeks, we keep the channel; below that, it goes to review.” Deciding the threshold after you see the data is how agencies talk their way out of bad results.

How small is too small to detect a lift at all?

The platforms will tell you where their own floor is. Meta states that to run a Conversion Lift test, your ad account needs a campaign from the past year with $5,000 USD or more in spend and a minimum of 500 conversions, prorated upward for longer campaigns — a 180-day campaign requires $10,000 and 1,000 conversions (About Conversion Lift, Meta Business Help Center).

Read that as a warning, not a rule. A plumber doing 40 booked jobs a month is nowhere near 500 conversions, and never will be. At 40 jobs a month, a 10% true lift is 4 jobs — a number that vanishes inside one rainy week. You are not going to measure a 10% effect. You can still measure a 40% effect, because a 40% effect is 16 jobs and 16 jobs is visible from the truck.

Monthly booked jobsWhat a holdout can realistically detectWhat to do instead
Under 25Almost nothingSkip the test; track cost per booked job and cut on that
25-75Large swings only (roughly a third of your volume or more)Run the whole channel dark, not one campaign
75-200Meaningful channel-level effects over 8-12 weeksGeo holdout, whole channel, two metros
200+Channel and campaign-type effectsGeo holdout, then platform lift tools

The honest version of this advice is the one most agencies won't say out loud: at small volume you cannot prove incrementality, so stop pretending a dashboard does. Manage to cost per booked job, keep the ad account in your own name, and reserve the holdout for the moment a channel is big enough to matter.

What do you do when the test says the channel adds nothing?

You get three options, and “ignore it” is not one of them: cut the channel, rebuild it, or accept it as a defensive cost with eyes open. Which one depends on why the test failed, and there are only a few real reasons.

  • The channel is buying customers you already had (brand terms, retargeting a warm list) — cut it or shrink it to a defensive floor.
  • The channel works but the offer doesn't. The clicks arrived and nobody booked. That's a landing page and speed-to-lead problem, not a media problem — rebuild before you cut.
  • The test was contaminated: budget still leaking into the control metro, a promo that ran in one market only, a competitor that went dark at the same time. Re-run it clean.
  • The effect is real but smaller than your noise floor. You didn't learn 'no'. You learned 'not measurable at this size'.

A failed test is not a wasted quarter. It is the cheapest thing you'll buy all year, because it stops a five-figure annual spend that was billing you for revenue you were already earning.

Which channel should you test first?

Test branded search first, because it is the single most likely line item to be paying for sales you already had — that's the exact finding of the eBay experiments, where brand-keyword ads showed no measurable short-term benefit. It's also the cheapest test to run: pause brand terms in one metro for a month and watch whether total booked jobs from that metro move at all.

Test orderChannelWhy it's hereTypical test window
1Branded searchHighest cannibalization risk, cheapest to pause4-6 weeks
2Retargeting / remarketingOverwhelmingly touches people already in your funnel4-8 weeks
3Paid social prospectingReal reach, but attribution is the weakest of any channel8-12 weeks
4Non-brand searchUsually the genuinely incremental one — test it last, and be ready to be wrong8-12 weeks

Verdict: if you only ever run one incrementality test, run it on branded search. It's the fastest, cheapest test on the list and it's the one most likely to hand you money back. Non-brand search is usually the last thing you should touch — but “usually” is a hypothesis, and the point of a holdout is to stop guessing. We build this kind of measurement into our paid ads program rather than selling it as a separate audit line.

How does a holdout test support a 90-day kill switch?

A 90-day kill switch says any channel that hasn't produced qualified leads in 90 days gets cut — and a geo holdout is what turns that rule from an opinion into evidence. Without a control group, “the channel isn't working” is a shouting match between you and whoever manages the spend. With one, it's a number in a spreadsheet that both sides agreed to before the test started.

That's also why the contract matters. If you're locked into 12 months, a failing channel is somebody else's revenue and nobody has an incentive to prove it dead. Month-to-month means the test result is actionable in 30 days, not next February. Anyone who resists running the holdout is telling you what they expect it to show.

If you want a second pair of eyes on whether your ad spend is buying customers or renting them, get a free audit — we'll look at the account, the tracking, and whether a holdout is even statistically possible at your volume before anyone suggests you pay for one.

Where does this fit in your stack?

If you're running a US service business, the playbook in this post pairs with our full services lineup and applies cleanly across our supported industries and US locations. If you want help implementing it, book a free strategy call — we'll review your current setup and prioritize the next three moves.

For the deeper engagement details, see our paid ads service. New to the terminology here? Our SEO & marketing glossary defines every acronym in this post.

Want this built for your vertical? See SEO for HVAC Companies, SEO for Roofing Contractors, SEO for SaaS Startups.

What are the most common questions about this topic?

Common questions readers send us about this topic.

What is incrementality testing in marketing?

Incrementality testing measures how much revenue a channel actually causes, by comparing a group that sees the ads against a control group that doesn't. The difference between the two is the lift. It answers the question a dashboard can't: if you switched the channel off, would you lose those sales, or would those customers have found you anyway? It requires a control group, which is what separates it from every attribution report you've ever been sent.

How is incrementality different from attribution?

Attribution divides credit among clicks that already happened. Incrementality asks whether the sale would have happened without the ad. They can disagree wildly. In eBay's large-scale field experiments (NBER Working Paper 20171, 2014), brand-keyword ads had no measurable short-term benefit — yet every one of those clicks would appear as a conversion in an attribution report. Attribution is bookkeeping; incrementality is a causal experiment.

How do I run a geo holdout test on a small budget?

Pick two comparable metros, pull 8 to 12 weeks of baseline booked jobs for each, switch the channel completely off in one of them, and compare booked jobs per capita over the test window. Change nothing else during the test. The cost is not software; it's the revenue you might forgo in the dark market. Pick a control metro you can afford to lose ground in for two months.

How long should a holdout test run?

Plan for 6 to 12 weeks. The window has to cover at least one full sales cycle after you go dark, plus enough weeks that a normal bad month can't explain the difference. Home services with a one-week close can get a signal in 6 weeks. Anything with a 60-day sales cycle needs the pipeline to turn over once before the clock even starts. Two-week tests measure weather, not advertising.

Can a business doing 40 jobs a month measure incrementality?

Not for small effects. At 40 booked jobs a month, a 10% lift is 4 jobs, which disappears inside normal week-to-week variation. You can still detect a large effect — a 40% swing is 16 jobs and that's visible. For context, Meta requires $5,000 in spend and 500 conversions before it will even run a Conversion Lift test. Below that scale, manage to cost per booked job instead.

Does pausing brand ads count as an incrementality test?

It's the best cheap one available. Pause branded search in one metro for four to six weeks and watch whether total booked jobs from that metro fall. If organic and direct absorb the demand and bookings hold, the brand campaign was buying customers you already had. Keep a defensive floor if competitors bid on your name, but stop counting those conversions as growth.

What is a dark period in media testing?

A dark period is a defined block of time when you switch a channel completely off, so you can compare what happened against what your model predicted would have happened. It's the single-market fallback when you don't have two comparable metros. Google's open-source CausalImpact package builds that counterfactual with a Bayesian time-series model — but it warns that valid conclusions need control series the intervention didn't touch.

Do I need a data science team to test incrementality?

No. A two-metro geo holdout needs a spreadsheet, a defined start and end date, booked jobs per market, and the discipline to change nothing else. The math is a comparison of jobs per capita between treatment and control. Modeling tools help at scale, but the hard part is never the statistics — it's keeping the control market genuinely dark and writing the decision rule down before you look at the data.

About the author

Hyder Shah

Founder & CEO, Foundgrove

Hyder Shah is the founder of Foundgrove, an SEO and GEO agency for US service businesses. See our editorial policy for how these guides are researched and reviewed.

Want help applying this to your business?

Book a free 30-minute call. We'll review your current acquisition stack and show you the three highest-leverage moves for your industry and state. Or read how our paid ads service works.

Free SEO & AI visibility auditGet my free audit