QuQi
ANALYTICS & REPORTING

Incrementality holdout design

Use when a channel claims credit you doubt and you want to test it rather than argue about models.

incrementality-holdout-design.md
Download .md
You are designing a holdout test to measure incremental effect. You are not evaluating a finished test and you are not defending the current attribution model.

Activity to test: {{ACTIVITY_TO_TEST}}
Current reported performance and spend: {{CURRENT_PERFORMANCE}}
Units we can split on, such as regions, stores, or audiences: {{SPLITTABLE_UNITS}}
Baseline volume and its week to week variation: {{BASELINE_VARIATION}}

Output:
1. The design: what is held out, for how long, and how units from {{SPLITTABLE_UNITS}} are assigned to each group. Name the assignment method.
2. A power estimate: the smallest effect this design could detect given {{BASELINE_VARIATION}}, with the arithmetic shown, and how much longer it would need to run to halve that.
3. A table: Threat to validity | How the design handles it | Residual risk. Cover spillover between units, seasonality, other activity running at the same time, and anything in {{CURRENT_PERFORMANCE}} suggesting the baseline is already trending.
4. Decided before launch: primary metric, guardrails, and the exact comparison you will run at the end.
5. The cost of the holdout in forgone {{ACTIVITY_TO_TEST}} results, and what finding would justify paying it.

Rules:
- If the design cannot detect an effect small enough to matter, say so and stop rather than proposing a test that will end ambiguous.
- Do not present an attribution difference as an incrementality result.
- Name in advance what would make you stop the test early, and what would not.

Fill in before running

Replace each placeholder with your own detail. The more specific you are, the less the model invents.

  • {{ACTIVITY_TO_TEST}}
  • {{CURRENT_PERFORMANCE}}
  • {{SPLITTABLE_UNITS}}
  • {{BASELINE_VARIATION}}

Getting a better result

  1. Split on geography where you can; user level holdouts leak through shared devices and households.
  2. Agree the analysis with whoever owns the channel before launch, not after the result arrives.
  3. A test that only detects a 40 percent effect is not worth running on a channel you think does 10.

Questions about this prompt

When do I design a holdout rather than change attribution model?

When the argument is about whether a channel causes conversions at all. Attribution models redistribute credit between channels. A holdout measures what happens when the activity stops. Switching model produces new numbers and the same disagreement, which is why this prompt refuses to present an attribution difference as an incrementality result.

What do I need in front of me?

The activity to test, its current reported performance and spend, the units you can split on such as regions, stores or audiences, and baseline volume with its week to week variation. That variation figure drives the power estimate. Without it nobody can say what effect size the test could ever detect.

What comes back?

The design with its named assignment method, a power estimate giving the smallest detectable effect with the arithmetic shown, a validity threat table covering spillover, seasonality and concurrent activity, the metrics and comparison agreed before launch, and the cost of the holdout. The power estimate decides whether to run it at all.

What is the mistake that costs me here?

Running a test that can only detect a 40 percent effect on a channel you believe delivers 10. It ends ambiguous and the argument restarts with the same people. Split on geography where you can, since user level holdouts leak through shared devices, and agree the analysis with the channel owner before launch.