Test hypothesis writer
Use before building a test, when you have an idea but no properly stated hypothesis.
Fill in before running
Replace each placeholder with your own detail. The more specific you are, the less the model invents.
- {{IDEA}}
- {{EVIDENCE}}
- {{LOCATION}}
- {{PRIMARY_METRIC}}
- {{BASELINE_RATE}}
- {{WEEKLY_VOLUME}}
Getting a better result
- Give the real baseline rate, otherwise the sample size reasoning is theatre.
- If it grades the evidence as Weak, go and get the evidence rather than running the test.
- Write point 7 down and keep it - it is what stops a null result being written off as a failure.
Questions about this prompt
When do I write a hypothesis rather than just build the test?
Before anything is built, when you have an idea and a hunch rather than a statement somebody could disprove. It grades your evidence as Strong, Moderate or Weak and says the test should not be prioritised when it is Weak. That grade is the reason to run this, so do not skip past it to the wording.
What do I need in front of me before running it?
The idea, the evidence that prompted it, the page or flow, the primary metric, the real baseline rate and weekly volume on that page. Guessing the baseline makes the sample size reasoning theatre, because the minimum detectable effect is worked out from it. If nobody can find the baseline, that is already a measurement finding.
What comes back, and which part is worth acting on?
Seven fields: problem statement, evidence grade, the because and we believe hypothesis, a primary and a guardrail metric, a realistic minimum detectable effect with the reasoning shown, the reject condition, and what a null teaches you. Point five is where most ideas die, since it will say plainly when the traffic cannot carry a test.
What is the mistake that costs me here?
Writing point seven after the result arrives instead of before. Stating in advance what a null teaches you is what stops it being written off as a failed test later. If the evidence grades as Weak, go and get evidence rather than running anyway, and do not accept an uplift estimate the inputs cannot support.