Sign in Start free
CONVERSION

Post test readout including null results

Use when a test has finished and you need an honest write-up, including when nothing moved.

post-test-readout.md
Download .md
You are writing the readout for a completed experiment. Honesty about a null result is the point of this document.

TEST NAME: {{TEST_NAME}}
HYPOTHESIS AS WRITTEN BEFORE THE TEST: {{ORIGINAL_HYPOTHESIS}}
VARIANTS: {{VARIANT_DESCRIPTIONS}}
DATES AND DURATION: {{DATES}}
SAMPLE PER VARIANT: {{SAMPLE_SIZES}}
PRIMARY METRIC RESULTS: {{PRIMARY_RESULTS}}
GUARDRAIL AND SECONDARY METRICS: {{SECONDARY_RESULTS}}
ANYTHING UNUSUAL DURING THE TEST: {{ANOMALIES}}

Write the readout with these sections:
1. Result in one sentence - one of: variant won, variant lost, no detectable difference, or inconclusive due to insufficient power. Choose one and do not hedge across two.
2. The numbers, with confidence interval or significance as supplied. If I have not supplied significance, say it cannot be judged rather than calculating from incomplete inputs.
3. What this means for the hypothesis - was the underlying belief supported, contradicted, or left untested.
4. Guardrail check - anything that got worse.
5. Segments - only if I supplied segment data. If I did not, state that segment claims would be post hoc and are not included.
6. Decision - ship, revert, iterate, or investigate, with the reason.
7. What we learned even if nothing moved. A null result usually means the change was too small to matter or the problem is elsewhere. Say which you think it is and why.

Rules: never describe a non-significant difference as a "slight lift" or a "positive trend". Do not go hunting for a winning segment to rescue the result. Do not open with scene setting or close with a paragraph restating section one. No em dashes.

Fill in before running

Replace each placeholder with your own detail. The more specific you are, the less the model invents.

  • {{TEST_NAME}}
  • {{ORIGINAL_HYPOTHESIS}}
  • {{VARIANT_DESCRIPTIONS}}
  • {{DATES}}
  • {{SAMPLE_SIZES}}
  • {{PRIMARY_RESULTS}}
  • {{SECONDARY_RESULTS}}
  • {{ANOMALIES}}

Getting a better result

  1. Paste the hypothesis exactly as it was written beforehand, not a tidied version.
  2. If you did not pre-register segments, let it refuse to report them - that rule is doing real work.
  3. Section seven is what makes null results worth running; keep a running file of them.

Questions about this prompt

When do I write a full readout rather than share the numbers?

Whenever a test ends, and particularly when nothing moved. Writing up only the winners quietly builds a record that overstates your hit rate and hides what the losses taught you. The readout picks one verdict, won, lost, no detectable difference or underpowered, and will not hedge across two of them.

What do I need in front of me before running it?

The hypothesis exactly as written beforehand rather than a tidied version, the variants, dates, sample per variant, primary and guardrail results, and anything unusual during the run. If you did not supply significance it will say significance cannot be judged rather than calculating one from incomplete inputs, which is worth leaving in place.

What comes back, and which part is worth acting on?

Seven sections opening with the single sentence verdict. Section seven is the one to keep: whether a null means the change was too small to matter or the problem lies elsewhere. Section five refuses to report segments you did not supply in advance, and that refusal is doing more work than it appears to.

What is the mistake that costs me here?

Hunting for a segment that rescues the result. Slice enough ways and you will always find one, which is why the prompt declines to. The related habit is calling a non-significant difference a slight lift or a positive trend, which turns noise into a decision and eventually into a rollout nobody can account for.