The right protocol is not the most sophisticated; it creates a credible comparison at the level where the action can…
The right protocol is not the most sophisticated; it creates a credible comparison at the level where the action can be affected. Individual randomisation fits controllable exposure and weak interference. A geo test fits territorial activation. A quasi-experiment can document a non-random rollout, but its conclusion relies on stronger assumptions. Without comparable unit, prior measurement and decision rule, the result remains descriptive.
Decision and method.
Choose the simplest protocol that identifies the effect useful to the decision, and refuse any causal conclusion when comparison conditions cannot be verified. Start from the planned change, its unit of application and the threshold that would alter action; an available metric does not define an experiment. Identify the smallest assignment level without major leakage: individual, account, store, area or period. If stable random assignment is feasible, use a controlled test; otherwise document assignment rule and quasi-experimental assumptions. Set minimum useful effect, variability, unit count, duration and diffusion risk. Define population, metric, horizon, exclusions and continue/stop/deepen rule before reading results. Compare observed interval with decision threshold: imprecision calls for new measurement or a prudent decision, not artificial certainty.
Worked example.
In an illustrative case, activation is deployed in 12 test areas and compared with 12 control areas. Average weekly revenue before test is €200,000 per group. Observed mean weekly difference is €12,000, with weekly interval from −€4,000 to €28,000. The point estimate is 12,000 ÷ 200,000 = 6%. The interval crosses zero and includes very different decision effects; the central value cannot authorise national generalisation. The result is indeterminate: extend or redesign the test if the minimum useful effect justifies cost, rather than scale activation from the point estimate alone.
Checks and limits.
Assignment unit matches the controllable lever; comparison, exclusions and metric are set before outcomes; contamination and simultaneous changes are documented; decision uses an interval and business threshold, never the point estimate alone. A local experiment does not guarantee the same effect in another population or period; platforms can link groups through learning mechanisms; a quasi-experiment relies on partly untestable assumptions; no detected difference does not prove no effect at all.
Resources and sources.
Download the protocol selector. Related: define incrementality, design a geo test, triangulate evidence, compare measurement families. Sources: Google Ads experiments; Google Ads Performance Max incrementality; Liu, Bettaney & Chamberlain (2018); Callaway & Sant’Anna; Abadie.
