QuQi

Designing An SEO Test With A Holdout

Before-and-after on a page group is not evidence, because seasonality, core updates and other releases move the treated pages and everything else together. The obvious correction, holding back a few pages as a control, fails when the control is picked by hand and ends up systematically different from the treated set. This applies when a template-level change can be applied to some URLs and withheld from others at a scale of a few hundred pages; below that, accept you are making a judgement call and say so rather than dressing it up as a test.

Obtenir le fichier de compétence Laissez les agents s’en charger
CATÉGORIE
Analyse et rapports
FORMAT
organic-holdout-test-design.md
ÉTAPES
7
PRIX
Gratuit — sans compte
QUAND S’EN SERVIR

Use when someone wants proof that an SEO change caused a result and a before-and-after chart will not settle the argument.

Le fichier de compétence

organic-holdout-test-design.md
---
name: organic-holdout-test-design
description: Use when someone wants proof that an SEO change caused a result and a before-and-after chart will not settle the argument.
---

# Designing An SEO Test With A Holdout

Before-and-after on a page group is not evidence, because seasonality, core updates and other releases move the treated pages and everything else together. The obvious correction, holding back a few pages as a control, fails when the control is picked by hand and ends up systematically different from the treated set. This applies when a template-level change can be applied to some URLs and withheld from others at a scale of a few hundred pages; below that, accept you are making a judgement call and say so rather than dressing it up as a test.

## Ce qu’il vous faut d’abord

- a change that can be applied to some URLs and withheld from others without breaking the site
- at least a few hundred comparable URLs, ideally from one template
- daily Search Console clicks and impressions per URL for 8 weeks before the test
- the test length and the single success metric, written down before launch

## Méthode

1. Set the unit of assignment as the URL and check the candidate URLs are genuinely comparable: same template, similar impression range, similar intent. A mixed set produces a control group that drifts for reasons of its own.
2. Randomise assignment rather than choosing control pages by hand, then verify balance on the 8 weeks of pre-period impressions. If the two arms already differ before you change anything, reshuffle and check again.
3. Calculate the minimum detectable effect from the pre-period variance before launch. Many single-template tests cannot detect anything under roughly 10 percent, and knowing that in advance stops you spending six weeks to learn nothing.
4. Ship the change to the treated arm on one day, and ship nothing else that touches only one arm. A release across both arms is survivable; a release across one ends the test.
5. Start the measurement window at recrawl, not at deploy. Counting from the deploy date mixes in days when the change was not yet in the index and drags the measured effect towards zero.
6. Compare treated against control as a ratio over time rather than as two separate before-and-after numbers. The ratio absorbs seasonality and update volatility, which is the whole reason the control exists.
7. Report the effect as an interval and publish the result even when it is null. Extending the test until the lines separate turns it into a search for a favourable week.

## Ce que ça produit

A two-page test record covering the randomisation method, the pre-period balance check, the minimum detectable effect and a result stated as an interval.

## Là où ça dérape

- Using a time-based control, the same pages before and after, which cannot separate your change from anything else that happened that month
- Stopping the moment the two lines separate, which catches noise at its widest and reports it as a win
- Testing on the pages you most want to fix, so the treated arm is a set of outliers and the result does not carry to the rest of the template
- Running two tests over overlapping URL sets, after which neither result can be attributed to either change

---

Extrait de la bibliothèque de compétences QuQi - https://www.quqi.io/fr/skills/organic-holdout-test-design
Téléchargement gratuit · sans compte, sans e-mail

Ce qu’il vous faut d’abord

  • a change that can be applied to some URLs and withheld from others without breaking the site
  • at least a few hundred comparable URLs, ideally from one template
  • daily Search Console clicks and impressions per URL for 8 weeks before the test
  • the test length and the single success metric, written down before launch

Méthode

  1. 01 Set the unit of assignment as the URL and check the candidate URLs are genuinely comparable: same template, similar impression range, similar intent. A mixed set produces a control group that drifts for reasons of its own.
  2. 02 Randomise assignment rather than choosing control pages by hand, then verify balance on the 8 weeks of pre-period impressions. If the two arms already differ before you change anything, reshuffle and check again.
  3. 03 Calculate the minimum detectable effect from the pre-period variance before launch. Many single-template tests cannot detect anything under roughly 10 percent, and knowing that in advance stops you spending six weeks to learn nothing.
  4. 04 Ship the change to the treated arm on one day, and ship nothing else that touches only one arm. A release across both arms is survivable; a release across one ends the test.
  5. 05 Start the measurement window at recrawl, not at deploy. Counting from the deploy date mixes in days when the change was not yet in the index and drags the measured effect towards zero.
  6. 06 Compare treated against control as a ratio over time rather than as two separate before-and-after numbers. The ratio absorbs seasonality and update volatility, which is the whole reason the control exists.
  7. 07 Report the effect as an interval and publish the result even when it is null. Extending the test until the lines separate turns it into a search for a favourable week.

Ce que ça produit

A two-page test record covering the randomisation method, the pre-period balance check, the minimum detectable effect and a result stated as an interval.

Là où ça dérape

  • Using a time-based control, the same pages before and after, which cannot separate your change from anything else that happened that month
  • Stopping the moment the two lines separate, which catches noise at its widest and reports it as a win
  • Testing on the pages you most want to fix, so the treated arm is a set of outliers and the result does not carry to the rest of the template
  • Running two tests over overlapping URL sets, after which neither result can be attributed to either change

Utiliser cette compétence dans votre propre IA

Le fichier téléchargé est un simple markdown dont l'en-tête porte le nom et le déclencheur. Quand un assistant sait charger des compétences tout seul, c'est cet en-tête qu'il lit pour décider que celle-ci s'applique.

Claude Code Enregistrez-le sous ~/.claude/skills/organic-holdout-test-design/SKILL.md et Claude le charge tout seul dès que ce que vous faites correspond au déclencheur. Placez-le plutôt dans .claude/skills d'un projet si toute l'équipe doit l'avoir.
Claude Importez le fichier dans la section compétences de vos réglages. Une fois là, il s'applique tout seul dans toute conversation où le déclencheur colle, sans que vous ayez à y penser.
ChatGPT Il n'existe pas de format de compétences où l'installer, alors collez le contenu du fichier dans les instructions d'un Projet ou d'un GPT personnalisé. Il s'applique ensuite à toutes les conversations du projet, pas seulement à celle où vous l'avez collé.
Tout le reste Collez le markdown dans la conversation avant votre question. Cela fonctionne avec n'importe quel assistant, il faut simplement le recoller à chaque fois.

Plus dans Analyse et rapports