Sign in Start free
ANALYTICS & REPORTING

A/B test result reading

Use when a test has finished and you need to know whether the result means anything.

ab-test-result-reading.md
Download .md
You are reviewing an A/B test result critically before anyone ships the winner.

Test: {{TEST_DESCRIPTION}}
Hypothesis as written before the test: {{HYPOTHESIS}}
Variant results: {{RESULTS_WITH_SAMPLE_SIZES}}
Runtime and traffic split: {{RUNTIME_AND_SPLIT}}
Primary metric declared in advance: {{PRIMARY_METRIC}}
Guardrail metrics: {{GUARDRAILS}}

Answer:
1. Was the test valid? Check sample size, runtime against full business cycles, split balance, sample ratio mismatch, and whether the primary metric was set before the test.
2. What does the primary metric result actually say? Give the observed difference, the uncertainty, and a plain sentence a non-analyst would understand.
3. Guardrails: did anything get worse? Say so even if the primary metric won.
4. Ship, do not ship, or run longer. One of those three, with the reason.
5. If someone points at a segment where the result looks stronger, explain why that is or is not evidence.

Rules:
- Do not call a result significant without stating the sample sizes and the metric it applies to.
- Treat any metric not declared in advance as exploratory and label it so.
- If the test ran for less than one full weekly cycle, say the result is unreliable regardless of the numbers.
- Never recommend shipping on a directional result alone. Say what it would cost to get certainty.

Fill in before running

Replace each placeholder with your own detail. The more specific you are, the less the model invents.

  • {{TEST_DESCRIPTION}}
  • {{HYPOTHESIS}}
  • {{RESULTS_WITH_SAMPLE_SIZES}}
  • {{RUNTIME_AND_SPLIT}}
  • {{PRIMARY_METRIC}}
  • {{GUARDRAILS}}

Getting a better result

  1. Paste the pre-test hypothesis verbatim - rewriting it afterwards is the most common way tests lie.
  2. Include guardrails even if they look fine; a conversion win with a refund rise is not a win.
  3. Ask what the minimum detectable effect was, and whether the test could ever have found the effect claimed.

Questions about this prompt

When do I use this rather than reading the testing tool verdict?

Before shipping a winner, especially when the result is close or someone has found a segment where it looks better. Testing tools report significance on whatever metric you point at. They do not check whether the runtime covered a full business cycle, or whether the primary metric was chosen before the test ran.

What do I need in front of me?

The pre-test hypothesis copied verbatim, variant results with sample sizes, runtime and traffic split, the primary metric declared in advance, and the guardrails. Paste the hypothesis as written, not as you remember it. Quietly rewriting it after seeing the result is the most common way a test tells you something false.

What comes back?

A validity check covering sample size, runtime against business cycles, split balance and sample ratio mismatch, what the primary metric actually says in plain words, guardrail movement, then one of ship, do not ship or run longer, and why a strong looking segment is or is not evidence. That three way verdict is the useful bit.

What is the mistake that costs me here?

Shipping on a directional result because the deadline arrived. The prompt will not recommend it, and instead tells you what certainty would cost, which is the figure to take to whoever is pushing. Ask for the minimum detectable effect too: a test underpowered for the effect claimed could never have found it.