How to run a geo holdout that anyone will believe
The method is not difficult. What makes it fail is almost always a decision taken before the test started.
iStudios, Measurement practice
A geo holdout is the most useful measurement instrument most advertisers never use. You switch a channel off in a set of regions, leave it on in a comparable set, and compare what happens. It produces a causal estimate, which platform reporting cannot, and it works whatever happens to cookies.
It is also the test people most often run badly, in ways that are obvious afterwards and invisible at the time.
Decide the question before the design
"Does channel X work" is not a question a test can answer. "If we removed channel X at current spend levels for eight weeks, how much revenue would we lose" is. The second version commits you to a spend level, a duration and an outcome measure, which is precisely why people prefer the first.
Match the regions on outcome, not on population
Test and control groups should be matched on the historical behaviour of the thing you are measuring, not on how many people live there. Take twelve months of weekly revenue by region, and pair regions whose series move together. Similar level, similar seasonality, similar response to past promotions.
If test and control did not track each other before the test, nothing they do during it means anything.
Size it honestly, then decide whether to proceed
- Work out the smallest effect that would change a decision. If a 3% lift and a 6% lift lead to the same action, you only need to detect 6%.
- Calculate the duration needed to detect it given your weekly variance. Do this before committing, not after.
- If the honest answer is thirty weeks, you have learned something important: this channel cannot be tested at this spend level. Change the spend or change the question.
- Never shorten a test because the early numbers look good. Early numbers always look like something.
Protect the test from the organisation
The most common cause of failure is not statistical. It is a regional sales manager noticing their area has gone quiet and compensating with a local promotion, or a platform's automated bidding reallocating budget into the control regions because that is where conversions now are. Freeze what you can, monitor what you cannot, and document every intervention.
Publish the result you got
Including the confidence interval, including the inconvenient direction, including the caveats. A measurement practice earns its credibility by reporting the tests that went against its own recommendation, and it spends that credibility every time it quietly reruns one.