When geo-lift testing is useful
A geo-lift test changes marketing activity in selected regions and compares the resulting outcome with untreated regions. It is especially useful when user-level randomisation is unavailable, cross-device identity is incomplete, privacy limits path-level data, or the channel operates at a market level.
Typical questions include whether a channel produces incremental revenue, whether an upper-funnel campaign creates demand, and whether a regional budget increase earns more than its cost. Geo testing can evaluate online or offline outcomes as long as they are measured consistently by geography and time.
The model can quantify uncertainty, but it cannot repair a contaminated control group or an intervention too small to detect.
Minimum data requirements
OpenLift’s minimum modelling dataset contains a date, geography, and outcome. Daily or weekly observations can be used. The outcome should be stable, economically meaningful, and measured in the same way across test and control markets.
Additional fields make the final decision more useful:
- Spend enables incremental ROAS, CAC, and profit calculations.
- Treatment and period flags make the test definition explicit.
- Channel and campaign preserve the intervention context.
- Holiday, weather, or event covariates can explain predictable external variation when supported by the model.
The pre-period should contain enough history to learn regular seasonality and the relationship between markets. More data does not automatically improve validity: a structural break or tracking change can make older history misleading.
How control-market matching works
Candidate control markets should behave like the test market before treatment without being exposed to the intervention. OpenLift ranks candidates using time-series similarity methods including dynamic time warping, Euclidean distance, and correlation distance.
Each metric captures something different. Correlation focuses on co-movement, Euclidean distance penalises level differences, and dynamic time warping tolerates some timing displacement. A high-ranking match is a design aid—not proof that the causal assumptions hold.
What to inspect before accepting a match
- Visual similarity of pre-period levels, trends, and seasonality.
- Stability of the relationship across multiple pre-period windows.
- Comparable business conditions, distribution, pricing, and market maturity.
- No campaign spillover, shared media footprint, or strategic change in controls.
- Residual behaviour that does not show a persistent pattern before treatment.
Power, MDE, and treatment duration
Power analysis asks whether the design can reliably detect the effect that matters. The minimum detectable effect (MDE) is the smallest lift the design can distinguish given the outcome variance, pre-period fit, number of markets, treatment duration, and decision threshold.
A test is not improved by declaring a smaller MDE after launch. If the planned effect is below the design’s sensitivity, change the intervention size, add suitable markets, improve the outcome signal, or extend the treatment window before spending.
| DESIGN LEVER | LIKELY EFFECT | TRADE-OFF |
|---|---|---|
| Longer treatment | More post-period signal | Higher cost and greater contamination risk |
| Stronger treatment | Larger expected effect | Operational and budget constraints |
| Better control fit | Lower counterfactual error | May reduce the eligible market pool |
| More stable outcome | Lower variance | May move away from the ultimate business KPI |
A defensible geo-test workflow
- Define the decision. State what will change if the result is positive, weak, or negative.
- Choose the outcome and unit. Confirm consistent geo-level measurement and economic relevance.
- Screen candidate markets. Exclude spillover, operational differences, and known concurrent interventions.
- Match and validate controls. Combine similarity scores with visual, contextual, and residual diagnostics.
- Plan power and duration. Set the MDE and treatment needed before launch.
- Freeze the specification. Record markets, windows, exclusions, covariates, and decision thresholds.
- Run and monitor. Watch for tracking failures and contamination without repeatedly peeking for significance.
- Estimate and decide. Review lift, uncertainty, economics, limitations, and the next experiment.
Critical assumptions and failure modes
The untreated relationship between test and control markets must remain stable through the post-period. Treatment should not affect control geographies. No unmodelled intervention should selectively move the test market at the same time. These assumptions are not fully testable from the outcome series alone, so domain knowledge remains part of the design.
Results can be weakened by short windows, sparse conversions, volatile revenue, poor matches, overlapping campaigns, pricing or distribution changes, and tracking breaks. Report these limitations with the result rather than hiding them behind a lift estimate.
