A geographic test compares treated areas with areas that represent their trajectory without the campaign.

A geographic test compares treated areas with areas that represent their trajectory without the campaign. Credibility depends less on the raw number of regions than on pre-test similarity, stable treatment, absence of contamination and enough power to detect the decision-relevant effect.

Decision and method.

Launch only when geography supports a stable counterfactual and the minimum detectable effect is below the gain required to act. Define outcome, horizon, population and the minimum uplift that repays campaign cost. Match markets on level, trend, seasonality, distribution, price and competitive mix; reserve atypical periods for robustness. Randomise treatment within pairs, freeze calendar, control other levers and monitor exposure leakage. Estimate treated-control difference, interval, pre-trends and placebos, plus cost per incremental unit and influential markets.

Worked example.

Historical within-pair standard deviation is 4.8%, with 12 pairs, two-sided 5% risk and 80% target power. Minimum detectable effect ≈ (1.96 + 0.84) × 4.8% ÷ √12 = 3.88%. Expected minimum economic gain is 3.2%. The test is slightly underpowered for its economic threshold: add pairs, extend the period or reduce variance with pre-test covariates.

Checks and limits.

Define economic threshold before power; require stable levels and pre-trends, documented randomisation, and monitoring of leakage, distribution shifts and local events. Few markets make inference sensitive to an influential pair; results may not transfer to major cities or untested zones; national campaigns, untargeted media and travel can contaminate control.

Resources and sources.

Download the matched-market dataset. Related: formalise the counterfactual, interpret incrementality, calibrate a model with a test, compare protocols. Sources: Vaver & Koehler (2011); Kerman, Wang & Vaver (2017); Imbens & Rubin (2015).