Scale Up / Scale Down availability: Creating new Scale Up / Scale Down tests is currently unavailable in the test setup flow. This article explains how test design affects result interpretation. For current availability information, see Scale Up / Scale Down Testing.
TL;DR
Read the estimated impact, efficiency, probability of direction, and credible interval together. Start by identifying what the experiment changed: a holdout test measures the contribution of advertising, while a budget-change test measures the effect of increasing or decreasing spend relative to baseline.
A useful result is precise enough to inform the decision you tested. Positive lift does not automatically mean profitable growth, and a near-zero estimate does not automatically mean spend can be reduced without risk.
Overview
Incrementality tests estimate how revenue or acquisitions changed because of a controlled marketing intervention.
This article helps marketers and analysts interpret completed experiments, distinguish useful evidence from uncertain results, and decide what to investigate before changing spend. It focuses on GeoLift reporting. The same principles of scope and uncertainty apply when reviewing Meta Conversion Lift, but the reporting experience may differ.
Wait until the experiment has completed its planned measurement and cooldown periods and its results are ready. An experiment marked Under Review does not yet have complete results.
Where to find your results
Open the test from the tests list. Results are organized into three tabs:
Results overview: your headline numbers, the attribution model comparison, and breakdowns by customer type and platform.
Deep dive: the revenue with versus without marketing chart, lift by period, and the marketing value distribution.
Configuration: the test setup and spend test tracker.
Review the headline estimates, uncertainty, available breakdowns, and supporting charts alongside the experiment’s configuration and spend delivery.
Before interpreting the numbers, confirm:
Which campaigns or advertising activity were tested.
Whether the experiment measured advertising contribution or a budget change.
The primary metric.
The test and comparison groups.
The planned spending levels and experiment dates.
Key terms
Marketing contribution: The estimated revenue effect of the tested marketing intervention, relative to the experiment’s comparison baseline.
Revenue lift: The estimated percentage change in revenue relative to the experiment’s comparison baseline. A 15% lift means revenue was estimated to be 15% higher than that baseline.
Incremental return on ad spend (iROAS): Incremental revenue divided by the spend included in the experiment’s analysis. Confirm the spend basis before comparing different test designs.
Probability of direction: The model’s probability that the estimated effect has the direction indicated in the results. Read the displayed label to understand which direction it describes.
Credible interval: A range of plausible values for the effect, given the data and model assumptions. A 90% credible interval contains 90% of the modeled probability. An HDI, or highest density interval, is a type of credible interval.
Synthetic control: A modeled estimate of what would have happened in the tested regions without the intervention, using comparison regions and baseline behavior.
How it works
GeoLift uses comparable geographic regions to estimate what changed because of the tested advertising intervention.
The comparison depends on the experiment’s design.
Holdout tests
In a holdout test for active advertising, the selected campaigns continue running in one group of regions and are withheld or reduced in another.
The analysis estimates the contribution of the advertising under the conditions tested.
Scale Up / Scale Down tests
A budget-change test compares a planned increase or decrease in spend with a matched group that remains at baseline.
The analysis evaluates the effect of that specific spend change. It does not automatically measure the contribution of the entire campaign.
For example, revenue declining after a budget reduction can indicate that the removed spend was contributing to results. That interpretation differs from a negative contribution estimate in a holdout test.
Why the comparison matters
The analysis uses the geographic comparison and baseline behavior to estimate what would likely have happened without the intervention.
A simple before-and-after comparison cannot isolate the effect. Revenue may also change because of seasonality, promotions, inventory, or other business conditions.
Read the estimated impact together with the uncertainty and the experiment’s delivery. The lift number alone is not enough.
How to read each metric
Marketing contribution: How large was the estimated effect?
Marketing contribution expresses the estimated revenue effect in dollars.
Consider whether the amount is meaningful for your business and whether the result range supports the same conclusion as the headline estimate.
For a holdout test, the estimate concerns the contribution of the tested advertising. For a budget-change test, interpret the effect relative to the spending change and comparison baseline.
For acquisition-based experiments, review the corresponding incremental-conversion and acquisition-lift metrics.
Revenue lift: How large was the effect relative to baseline?
Revenue lift expresses the estimated effect as a percentage.
Read it alongside the dollar or conversion impact. A large percentage change on a small baseline may represent less business value than a smaller percentage change on a larger baseline.
Keep the comparison specific to the tested regions, activity, primary metric, and dates.
iROAS: Did the incremental value justify the analyzed spend?
iROAS helps you evaluate the incremental revenue generated relative to the spend included in the analysis.
Before comparing it with another result, confirm that both use a comparable:
Revenue definition.
Campaign or channel scope.
Measurement period.
Spend basis.
Do not assume that total campaign spend and the additional or removed spend in a budget-change test are interchangeable denominators.
iROAS is a revenue-efficiency metric. It does not establish profitability on its own. Consider your margins, acquisition economics, and other business constraints.
Probability of direction: How strongly does the model support the direction?
Probability of direction describes how strongly the model supports the indicated positive or negative effect.
A high probability provides stronger evidence for that direction under the model’s assumptions. It does not tell you whether the effect is large enough to matter or whether the investment is profitable.
A displayed value close to 100% is not a guarantee. Read it alongside the credible interval and the quality of the experiment.
Credible interval: What outcomes remain plausible?
The credible interval shows the uncertainty around the estimate.
A narrower interval generally indicates a more precise estimate. A wider interval leaves more uncertainty about the size, and sometimes the direction, of the effect.
Ask whether the range includes outcomes that would lead to different business decisions.
For example, if a Scale Down test’s range includes both little change and a revenue loss your business could not accept, the test has not established that the reduction is low risk.
An interval crossing zero does not always make the result useless. A narrow range around zero can be informative if it rules out effects that would matter to the decision. A wide range that includes meaningful gains and losses is much harder to act on.
Compare against your attribution models
Where available, compare the GeoLift result with attribution models such as First Click, Last Click, Linear All, Linear Paid, and Triple Attribution.
Attribution assigns credit across customer touchpoints. GeoLift estimates the effect of a controlled advertising intervention.
A higher attributed ROAS than experimental iROAS may suggest that attribution is crediting purchases that would have happened anyway. However, the numerical gap is not automatically a direct measure of over-attribution.
Before drawing that conclusion, review:
The campaigns and channels included.
The outcome and revenue definitions.
The measurement periods and attribution windows.
The spend included in each calculation.
The experiment’s uncertainty and delivery.
Use the comparison to investigate differences. Do not treat the experimental result as an exact, permanent correction factor for an attribution model.
Read the breakdowns
Where available, review breakdowns by customer type or platform to understand how the estimated impact varies within the tested activity.
Review the evidence supporting each breakdown before using it to shift budget. A clear overall result does not guarantee equally strong evidence for every smaller segment. If several campaigns or channels were tested together, do not assume the combined result establishes each component’s independent effect.
Customer type: incremental impact split into New and Returning customers, each with its own iROAS and incremental revenue. This tells you whether the channel is acquiring new customers or mostly driving repeat purchases.
Platform: incremental impact split by platform or channel included in the test.
Use these to decide not just whether to scale, but where the scalable value actually is. A channel that is strongly incremental for new customers is a different decision than one that mostly lifts returning buyers.
What a decision-ready result looks like
A result is useful when the experiment ran reliably and the estimated effect is precise enough to answer your business question.
The desired outcome depends on the test.
For a holdout test
A meaningful positive contribution, strong evidence for a positive direction, and an interval that remains positive support the conclusion that the tested advertising contributed to results.
You still need to evaluate whether the incremental return justified the cost.
For a Scale Up test
Look for evidence that the additional spend generated enough additional value to meet your business goals.
A positive revenue effect alone does not establish that the increase was efficient. A successful increase at one spending level also does not guarantee the same return from the next increase.
For a Scale Down test
Look at the business impact of the reduction and the uncertainty around it.
A small measured decline may be acceptable if the savings justify it. Little or no measured decline may support a reduction if the result range rules out a loss your business would consider unacceptable.
The threshold for an acceptable loss should come from your business goals, not from whether the headline estimate is close to zero.
What an inconclusive result means, and what to do
A result is inconclusive for your decision when the evidence leaves materially different outcomes plausible.
For example, a wide interval may include both a worthwhile gain and a substantial loss. The experiment has not provided a clear enough basis for choosing between the corresponding budget decisions.
This does not establish that the advertising or budget change had no effect.
Before repeating the experiment:
Review the test design. Check for promotions, unplanned campaign changes, geographic overlap, or platform issues during the measurement period.
Review actual delivery. Confirm whether spend and campaign delivery followed the plan.
Check the original feasibility assessment. A test with limited power may be unable to distinguish the expected effect from normal variation.
Discuss the next measurement step with your Triple Whale team. A longer, larger, or cleaner follow-up test may provide a clearer answer.
Marketing mix modeling can provide additional planning context between experiments, but it does not turn an inconclusive experiment into a confirmed causal result.
How to read a negative result
Start by confirming what the experiment changed and what the displayed metric represents.
In a holdout test, a negative contribution estimate suggests that outcomes associated with the tested advertising were lower than the estimated comparison outcome. Review the uncertainty and test delivery before concluding that the advertising reduced performance.
In a Scale Down test, lower revenue or acquisitions relative to unchanged baseline spend can indicate that the removed spend was contributing to results. That can be a valid finding. The decision is whether the savings justify the estimated loss.
Before acting, review:
The credible interval and probability of direction.
Spend delivery and Data Health.
Geographic matching and separation between groups.
Regional events or business changes that affected the groups differently.
Unplanned campaign or budget changes.
Confirm the displayed sign convention before interpreting a negative number.
Discuss an unexpected or uncertain result with your Triple Whale team before making a major budget change.
When to use these results, and their limits
Use incrementality results to inform decisions about the activity and conditions included in the experiment.
Keep the following scope in view:
Campaigns and channels.
Geographic regions or audiences.
Budget levels and the size of any planned change.
Primary metric.
Measurement dates.
Business conditions during the test.
A test provides evidence about those conditions. It does not permanently define a channel’s value at every spending level or in every season.
A positive contribution result does not establish unlimited room to scale. A near-zero result after a modest reduction does not establish that all spend can be removed.
Use incrementality alongside attribution and MMM. Attribution supports ongoing campaign analysis, MMM provides broader planning context, and experiments validate specific causal questions over defined periods.
