Post test readout including null results
Use when a test has finished and you need an honest write-up, including when nothing moved.
Fill in before running
Replace each placeholder with your own detail. The more specific you are, the less the model invents.
- {{TEST_NAME}}
- {{ORIGINAL_HYPOTHESIS}}
- {{VARIANT_DESCRIPTIONS}}
- {{DATES}}
- {{SAMPLE_SIZES}}
- {{PRIMARY_RESULTS}}
- {{SECONDARY_RESULTS}}
- {{ANOMALIES}}
Getting a better result
- Paste the hypothesis exactly as it was written beforehand, not a tidied version.
- If you did not pre-register segments, let it refuse to report them - that rule is doing real work.
- Section seven is what makes null results worth running; keep a running file of them.
Questions about this prompt
When do I write a full readout rather than share the numbers?
Whenever a test ends, and particularly when nothing moved. Writing up only the winners quietly builds a record that overstates your hit rate and hides what the losses taught you. The readout picks one verdict, won, lost, no detectable difference or underpowered, and will not hedge across two of them.
What do I need in front of me before running it?
The hypothesis exactly as written beforehand rather than a tidied version, the variants, dates, sample per variant, primary and guardrail results, and anything unusual during the run. If you did not supply significance it will say significance cannot be judged rather than calculating one from incomplete inputs, which is worth leaving in place.
What comes back, and which part is worth acting on?
Seven sections opening with the single sentence verdict. Section seven is the one to keep: whether a null means the change was too small to matter or the problem lies elsewhere. Section five refuses to report segments you did not supply in advance, and that refusal is doing more work than it appears to.
What is the mistake that costs me here?
Hunting for a segment that rescues the result. Slice enough ways and you will always find one, which is why the prompt declines to. The related habit is calling a non-significant difference a slight lift or a positive trend, which turns noise into a decision and eventually into a rollout nobody can account for.