Scale Up / Scale Down availability: Creating new Scale Up / Scale Down tests is currently unavailable in the test setup flow. This article explains how the testing approach differs from a geo holdout. See Scale Up / Scale Down Testing for more information.
TL;DR
GeoLift measures incrementality by comparing outcomes across geographic regions. In a holdout test, the selected advertising continues in one group and is paused or reduced in another. In a Scale Up / Scale Down test, spend changes in the test group while a matched comparison group remains at baseline.
The analysis accounts for baseline differences between regions to estimate the effect of the tested advertising intervention
Overview
This article explains how a GeoLift test works, so you understand what the test is doing before you read its results. It covers why Triple Whale tests by geography, how test and control regions work, how regions are selected, and how a test's required duration is determined. This article is written for marketers and analysts planning or reviewing an incrementality test.
A geo holdout is one GeoLift test design. Scale Up / Scale Down uses a different comparison: a planned budget change against spend that remains at baseline.
Key terms
GeoLift: Triple Whale's geography-based incrementality test, found under Incrementality, New Lift Test, Test Design.
Test group: The regions where the advertising activity being evaluated runs. In a holdout test, the selected advertising continues in these regions. In a Scale Up / Scale Down test, these regions receive the planned budget change.
Comparison group: Comparable regions used to estimate what would have happened without the tested intervention. In a holdout test, the selected advertising is paused or reduced in these regions. In a Scale Up / Scale Down test, spend remains at baseline.
Synthetic control: A model of what the test regions would have done without the change, built from the control regions' behavior.
Marketing contribution: The revenue your advertising actually caused, measured as the gap between actual results and the synthetic control.
Feasibility report: A pre-launch estimate of whether the test can detect an effect at your budget level.
How it works
Holdout tests
A holdout test splits comparable regions into two groups. The selected advertising continues in the test regions and is paused or reduced in the holdout regions.
Triple Whale builds a synthetic control: a modeled estimate of what would have happened without the tested advertising intervention, based on the comparison regions’ behavior. The analysis uses the difference between observed outcomes and that estimate to measure the advertising’s incremental contribution.
Scale Up / Scale Down tests
Scale Up / Scale Down tests evaluate a budget change for advertising that is already active.
Scale Up: Spend increases above baseline in the test regions.
Scale Down: Spend decreases below baseline in the test regions.
In both cases, a matched comparison group continues at baseline spend. The result estimates the effect of the budget change, rather than automatically measuring the contribution of the entire campaign.
For example, a decline in revenue after reducing spend can indicate that the removed spend was contributing to results. It should not automatically be interpreted as a problem with the experiment.
For more information, see Scale Up / Scale Down Testing: Measuring the Impact of Budget Changes.
How regions are selected. Regions are matched so the test and control groups behave similarly before the test begins. Good matching is what makes the comparison valid.
How long a test needs to run. Before launch, Triple Whale produces a feasibility report that estimates whether the test can detect an effect at your budget level. You are told this before the test starts, not after.
When to use / trade-offs
Incrementality experiments can compare audiences or geographic regions. A geographic test can compare advertising against a holdout or evaluate a planned spend change against regions that remain at baseline.
They are less affected by audience spillover, where users in the holdout see ads through other means.
They work across all channels simultaneously.
They do not depend on platform-specific holdout features.
💡 On inconclusive results: A test that comes back "inconclusive" is not a failure. It means the sample size was not large enough to detect the effect. This is valuable information: it tells you the test design needs adjustment, not that the channel has zero impact.
