Sign in Start free
CONVERSION

Test hypothesis writer

Use before building a test, when you have an idea but no properly stated hypothesis.

test-hypothesis-writer.md
Download .md
You are helping write a testable hypothesis. Do not accept a vague idea - force it into a falsifiable statement.

THE IDEA: {{IDEA}}
EVIDENCE THAT PROMPTED IT: {{EVIDENCE}}
PAGE OR FLOW: {{LOCATION}}
PRIMARY METRIC: {{PRIMARY_METRIC}}
CURRENT BASELINE RATE: {{BASELINE_RATE}}
WEEKLY TRAFFIC OR SESSIONS ON THAT PAGE: {{WEEKLY_VOLUME}}

Produce exactly these fields:
1. Problem statement - what we believe is going wrong, in behavioral terms, not design terms.
2. Evidence grade - Strong (quantitative and qualitative), Moderate (one source), or Weak (opinion). Say which.
3. Hypothesis - in the form: Because [evidence], we believe that [change] for [audience] will cause [effect on primary metric]. We will know this is true when [measurable outcome].
4. Primary metric and one guardrail metric that would tell us the change caused harm elsewhere.
5. Minimum detectable effect that is realistic at this traffic level, and roughly how long the test would need to run. Show the reasoning. If the traffic is too low to detect anything sensible, say so plainly and suggest a non-test alternative such as a painted door or a qualitative study.
6. What result would make us reject the hypothesis.
7. What we learn if it is null - state this before running, not after.

Rules: if the evidence is Weak, say the test should not be prioritized and explain why. Do not estimate uplift percentages you cannot support. No em dashes.

Fill in before running

Replace each placeholder with your own detail. The more specific you are, the less the model invents.

  • {{IDEA}}
  • {{EVIDENCE}}
  • {{LOCATION}}
  • {{PRIMARY_METRIC}}
  • {{BASELINE_RATE}}
  • {{WEEKLY_VOLUME}}

Getting a better result

  1. Give the real baseline rate, otherwise the sample size reasoning is theatre.
  2. If it grades the evidence as Weak, go and get the evidence rather than running the test.
  3. Write point 7 down and keep it - it is what stops a null result being written off as a failure.

Questions about this prompt

When do I write a hypothesis rather than just build the test?

Before anything is built, when you have an idea and a hunch rather than a statement somebody could disprove. It grades your evidence as Strong, Moderate or Weak and says the test should not be prioritised when it is Weak. That grade is the reason to run this, so do not skip past it to the wording.

What do I need in front of me before running it?

The idea, the evidence that prompted it, the page or flow, the primary metric, the real baseline rate and weekly volume on that page. Guessing the baseline makes the sample size reasoning theatre, because the minimum detectable effect is worked out from it. If nobody can find the baseline, that is already a measurement finding.

What comes back, and which part is worth acting on?

Seven fields: problem statement, evidence grade, the because and we believe hypothesis, a primary and a guardrail metric, a realistic minimum detectable effect with the reasoning shown, the reject condition, and what a null teaches you. Point five is where most ideas die, since it will say plainly when the traffic cannot carry a test.

What is the mistake that costs me here?

Writing point seven after the result arrives instead of before. Stating in advance what a null teaches you is what stops it being written off as a failed test later. If the evidence grades as Weak, go and get evidence rather than running anyway, and do not accept an uplift estimate the inputs cannot support.