QuQi

Designing An SEO Test With A Holdout

Before-and-after on a page group is not evidence, because seasonality, core updates and other releases move the treated pages and everything else together. The obvious correction, holding back a few pages as a control, fails when the control is picked by hand and ends up systematically different from the treated set. This applies when a template-level change can be applied to some URLs and withheld from others at a scale of a few hundred pages; below that, accept you are making a judgement call and say so rather than dressing it up as a test.

Get the skill file Let the agents run it
CATEGORY
Analytics & reporting
FORMAT
organic-holdout-test-design.md
STEPS
7
PRICE
Free - no account
WHEN TO REACH FOR THIS

Use when someone wants proof that an SEO change caused a result and a before-and-after chart will not settle the argument.

The skill file

organic-holdout-test-design.md
---
name: organic-holdout-test-design
description: Use when someone wants proof that an SEO change caused a result and a before-and-after chart will not settle the argument.
---

# Designing An SEO Test With A Holdout

Before-and-after on a page group is not evidence, because seasonality, core updates and other releases move the treated pages and everything else together. The obvious correction, holding back a few pages as a control, fails when the control is picked by hand and ends up systematically different from the treated set. This applies when a template-level change can be applied to some URLs and withheld from others at a scale of a few hundred pages; below that, accept you are making a judgement call and say so rather than dressing it up as a test.

## What you need first

- a change that can be applied to some URLs and withheld from others without breaking the site
- at least a few hundred comparable URLs, ideally from one template
- daily Search Console clicks and impressions per URL for 8 weeks before the test
- the test length and the single success metric, written down before launch

## Method

1. Set the unit of assignment as the URL and check the candidate URLs are genuinely comparable: same template, similar impression range, similar intent. A mixed set produces a control group that drifts for reasons of its own.
2. Randomise assignment rather than choosing control pages by hand, then verify balance on the 8 weeks of pre-period impressions. If the two arms already differ before you change anything, reshuffle and check again.
3. Calculate the minimum detectable effect from the pre-period variance before launch. Many single-template tests cannot detect anything under roughly 10 percent, and knowing that in advance stops you spending six weeks to learn nothing.
4. Ship the change to the treated arm on one day, and ship nothing else that touches only one arm. A release across both arms is survivable; a release across one ends the test.
5. Start the measurement window at recrawl, not at deploy. Counting from the deploy date mixes in days when the change was not yet in the index and drags the measured effect towards zero.
6. Compare treated against control as a ratio over time rather than as two separate before-and-after numbers. The ratio absorbs seasonality and update volatility, which is the whole reason the control exists.
7. Report the effect as an interval and publish the result even when it is null. Extending the test until the lines separate turns it into a search for a favourable week.

## What this produces

A two-page test record covering the randomisation method, the pre-period balance check, the minimum detectable effect and a result stated as an interval.

## Where this goes wrong

- Using a time-based control, the same pages before and after, which cannot separate your change from anything else that happened that month
- Stopping the moment the two lines separate, which catches noise at its widest and reports it as a win
- Testing on the pages you most want to fix, so the treated arm is a set of outliers and the result does not carry to the rest of the template
- Running two tests over overlapping URL sets, after which neither result can be attributed to either change

---

From the QuQi skill library - https://www.quqi.io/skills/organic-holdout-test-design
Free to download · no account, no email

What you need first

  • a change that can be applied to some URLs and withheld from others without breaking the site
  • at least a few hundred comparable URLs, ideally from one template
  • daily Search Console clicks and impressions per URL for 8 weeks before the test
  • the test length and the single success metric, written down before launch

Method

  1. 01 Set the unit of assignment as the URL and check the candidate URLs are genuinely comparable: same template, similar impression range, similar intent. A mixed set produces a control group that drifts for reasons of its own.
  2. 02 Randomise assignment rather than choosing control pages by hand, then verify balance on the 8 weeks of pre-period impressions. If the two arms already differ before you change anything, reshuffle and check again.
  3. 03 Calculate the minimum detectable effect from the pre-period variance before launch. Many single-template tests cannot detect anything under roughly 10 percent, and knowing that in advance stops you spending six weeks to learn nothing.
  4. 04 Ship the change to the treated arm on one day, and ship nothing else that touches only one arm. A release across both arms is survivable; a release across one ends the test.
  5. 05 Start the measurement window at recrawl, not at deploy. Counting from the deploy date mixes in days when the change was not yet in the index and drags the measured effect towards zero.
  6. 06 Compare treated against control as a ratio over time rather than as two separate before-and-after numbers. The ratio absorbs seasonality and update volatility, which is the whole reason the control exists.
  7. 07 Report the effect as an interval and publish the result even when it is null. Extending the test until the lines separate turns it into a search for a favourable week.

What this produces

A two-page test record covering the randomisation method, the pre-period balance check, the minimum detectable effect and a result stated as an interval.

Where this goes wrong

  • Using a time-based control, the same pages before and after, which cannot separate your change from anything else that happened that month
  • Stopping the moment the two lines separate, which catches noise at its widest and reports it as a win
  • Testing on the pages you most want to fix, so the treated arm is a set of outliers and the result does not carry to the rest of the template
  • Running two tests over overlapping URL sets, after which neither result can be attributed to either change

Use this skill in your own AI

The download is a plain markdown file with the name and trigger in its frontmatter. Where an assistant supports skills it can load itself, that frontmatter is what it reads to decide this one applies.

Claude Code Save it as ~/.claude/skills/organic-holdout-test-design/SKILL.md and Claude loads it on its own when what you are doing matches the trigger line. Put it in .claude/skills inside a project instead if the whole team should have it.
Claude Upload the file in the skills section of your settings. Once it is there it applies itself in any conversation where the trigger fits, so you do not have to remember it exists.
ChatGPT There is no skills format to install into, so paste the file contents into a Project instruction or a Custom GPT instead. It then applies to every chat in that project rather than only the one you paste it into.
Anything else Paste the markdown into the chat before your question. It works in any assistant, it just has to be pasted again each time.

Questions about this skill

When is a holdout test the right method rather than a before-and-after chart on the changed pages?

When someone wants proof a change caused a result, and the change can be applied to some URLs and withheld from others across a few hundred comparable pages. Before-and-after cannot separate your change from seasonality, a core update or another release. Below that scale, accept you are making a judgement call and say so rather than dressing it up as a test.

What do I need in hand before starting, and what happens if I start without it?

A change you can withhold without breaking the site, a few hundred comparable URLs from one template, eight weeks of daily per-URL clicks and impressions, and the success metric and test length written down before launch. The pre-period data is what yields a minimum detectable effect. Without it you can spend six weeks to learn the test could never have read anything.

What do I end up with, and which part of it actually gets used?

A short test record: the randomisation method, the pre-period balance check, the minimum detectable effect, and a result stated as an interval. The measurement that carries it is treated against control as a ratio over time, because that absorbs seasonality and update volatility, which is the whole reason the control exists. Start the window at recrawl, not at deploy.

What is the mistake that most often ruins this, and what does it cost?

Stopping the moment the two lines separate. That catches noise at its widest, reports it as a win, and is hard to walk back once it has circulated. Running the test on the pages you most want to fix costs the same way: the treated arm becomes a set of outliers and the result does not carry to the rest of the template.

More in Analytics & reporting