Skip to main content

Scale Up / Scale Down Testing: Measuring the Impact of Budget Changes

Understand how controlled budget changes help you evaluate where to invest more or reduce spend.

K
Written by Kassandra Villa Arroyo

Availability update: Creating new Scale Up / Scale Down tests is currently unavailable in the test setup flow. This article explains the testing approach. Contact your Triple Whale team to discuss your testing options.

Overview


Scale Up and Scale Down tests measure the incremental impact of changing spend on advertising that is already running. They help you evaluate whether additional budget generates enough additional revenue, or whether reducing spend causes a meaningful decline in results.

These tests compare a planned budget change with a matched geographic group that continues at its baseline spending level. The comparison helps separate the effect of the budget change from other changes in business performance.

Use the findings to guide budget decisions for the campaigns and subcategories included in the test. Results apply to the spending levels, geographic regions, and dates tested.

Key terms


  • Baseline spend: The business-as-usual spending level used as the reference for the test.

  • Test group: The geographic regions where the planned budget increase or decrease takes place.

  • Comparison group: Matched geographic regions where spend remains at baseline during the measurement period.

  • Ramp-up period: The period before measurement when both geographic groups run at their allocated baseline budgets.

The Two Test Types


Scale Up

A Scale Up test increases spend in the test group above its baseline while the comparison group continues at baseline.

The result estimates the additional revenue or acquisitions caused by the increase. This helps you evaluate whether the extra investment generates enough additional value to justify its cost.

Positive lift alone does not establish that the increase is profitable. Consider the size of the effect, the additional spend, the uncertainty in the result, and your business’s margins.

Scale Down

A Scale Down test reduces spend in the test group below its baseline while the comparison group continues at baseline.

The result estimates how revenue or acquisitions change when that spend is reduced. A measurable decline suggests that the removed spend was contributing to results. Little or no measured decline may support considering a reduction, provided the result is precise enough to rule out a meaningful loss.

A Scale Down test evaluates the specific reduction tested. It does not establish that the entire campaign could be paused without affecting performance.

Both test types compare against a baseline that's already running, so there's no need to launch anything new. You're measuring the marginal effect of a budget change on a campaign you're already investing in, with the campaign setup itself held constant. Only the weekly spend level differs between test cells.

Why This Matters for Budget Allocation


Campaigns and subcategories can respond differently to budget changes. A campaign that performs well at its current budget may generate less value from each additional dollar as spending increases.

Scale Up and Scale Down tests help you evaluate those differences:

  • Find where to invest more. A Scale Up test that shows strong incremental lift signals a campaign that can absorb additional budget efficiently.

  • Find where to pull back. A Scale Down test that shows little to no drop in results signals spend that could be reduced or reallocated with minimal downside.

  • Compare across subcategories. Running tests across multiple campaigns or subcategories builds a picture of where each incremental dollar works hardest, so budget can move toward the highest-return areas.

How a test is structured


New Scale Up / Scale Down test creation is currently unavailable. The explanation below describes the test design rather than an available self-service setup procedure.

Separate the geographic groups

The setup uses duplicated campaigns to separate advertising into two non-overlapping geographic groups, identified as Cell A and Cell B.

One group provides the baseline comparison. The other receives the planned budget change during the measurement period.

Although the test evaluates advertising that is already active, it still requires campaign setup to separate delivery between the geographic groups.

Allocate baseline budgets

The original business-as-usual budget is divided between the groups in proportion to what their geographies historically received.

This preserves an appropriate baseline for each group. It does not assume that both groups should receive equal budgets.

Run both groups at baseline during ramp-up

Both groups run at their allocated baseline budgets during the ramp-up period.

The ramp-up period is separate from the measurement period. The planned increase or decrease should not be treated as part of baseline delivery.

Apply the planned budget change

At the start of the measurement period, the test group receives the planned spend change. The comparison group continues at baseline.

Keep other campaign changes consistent with the agreed test design. Unplanned changes to budgets, geographic targeting, creative, or other campaign settings can make the result harder to interpret.

Understanding Your Results


Results estimate the effect of the budget change relative to what would likely have happened if spending had stayed at baseline. The analysis uses the comparison regions and baseline behavior; a simple before-and-after revenue comparison cannot isolate that effect.

Reading a Scale Up result

Additional revenue or acquisitions indicate that the increased spend contributed to results.

Before increasing investment further, consider

  • How much additional value the spend generated.

  • Whether that value justified the additional cost.

  • How much uncertainty remains in the estimate.

  • Whether the test delivered spend according to plan.

A successful increase at one spending level does not guarantee the same return from the next increase.

Reading a Scale Down result

A decline in revenue or acquisitions after reducing spend can indicate that the removed spend was contributing to performance. That outcome is not automatically a measurement problem.

Little or no measured decline can support considering a budget reduction. However, an estimate near zero is not enough on its own. If the result range still includes a meaningful loss, the test has not established that the reduction is low risk.

When the result is inconclusive

An inconclusive result means the test could not distinguish the effect of the budget change from normal variation with enough confidence.

It does not prove that the additional or removed spend had no impact.

Review the test’s spend delivery, duration, geographic design, and any unexpected business or campaign changes before deciding whether to repeat it.

Apply the result to the tested scope

Use each result to inform decisions about the campaigns, subcategories, budget change, geographic regions, primary metric, and dates included in that experiment.

Avoid treating one result as a permanent measure of a campaign’s value. Performance can change with spending levels, seasonality, creative, competition, and customer behavior.

Frequently Asked Questions


Can I create a Scale Up or Scale Down test now?

Creating new Scale Up / Scale Down tests is currently unavailable in the test setup flow. Contact your Triple Whale team to discuss your testing options.

Do I need to launch a new marketing strategy?

No. These tests evaluate a budget change for advertising that is already active.

However, the setup uses duplicated campaigns and separate geographic groups. Testing existing advertising does not mean the original campaign configuration remains untouched.

How are the geographic groups matched?

Matching uses historical spend and revenue patterns to identify geographic groups that behaved similarly before the planned budget change.

The comparison group helps estimate what would likely have happened in the test group if spending had remained at baseline.

Can I run Scale Up and Scale Down tests on the same campaign at the same time?

Discuss the design with your Triple Whale team before planning concurrent tests. Separate geographic assignments alone do not establish that the experiments will be independent.

The design needs to account for overlapping delivery, campaign budgets, and other changes that could affect the comparison.

How should I use these results to guide budget decisions?

Consider more investment when a Scale Up test shows that additional spend generated sufficient incremental value with enough confidence.

Consider a reduction when a Scale Down test provides sufficiently precise evidence that the tested reduction did not cause an unacceptable loss.

In both cases, review the findings alongside your margins, growth goals, inventory, and other business constraints.

Did this answer your question?