A/B test result reading
Use when a test has finished and you need to know whether the result means anything.
Fill in before running
Replace each placeholder with your own detail. The more specific you are, the less the model invents.
- {{TEST_DESCRIPTION}}
- {{HYPOTHESIS}}
- {{RESULTS_WITH_SAMPLE_SIZES}}
- {{RUNTIME_AND_SPLIT}}
- {{PRIMARY_METRIC}}
- {{GUARDRAILS}}
Getting a better result
- Paste the pre-test hypothesis verbatim - rewriting it afterwards is the most common way tests lie.
- Include guardrails even if they look fine; a conversion win with a refund rise is not a win.
- Ask what the minimum detectable effect was, and whether the test could ever have found the effect claimed.
Questions about this prompt
When do I use this rather than reading the testing tool verdict?
Before shipping a winner, especially when the result is close or someone has found a segment where it looks better. Testing tools report significance on whatever metric you point at. They do not check whether the runtime covered a full business cycle, or whether the primary metric was chosen before the test ran.
What do I need in front of me?
The pre-test hypothesis copied verbatim, variant results with sample sizes, runtime and traffic split, the primary metric declared in advance, and the guardrails. Paste the hypothesis as written, not as you remember it. Quietly rewriting it after seeing the result is the most common way a test tells you something false.
What comes back?
A validity check covering sample size, runtime against business cycles, split balance and sample ratio mismatch, what the primary metric actually says in plain words, guardrail movement, then one of ship, do not ship or run longer, and why a strong looking segment is or is not evidence. That three way verdict is the useful bit.
What is the mistake that costs me here?
Shipping on a directional result because the deadline arrived. The prompt will not recommend it, and instead tells you what certainty would cost, which is the figure to take to whoever is pushing. Ask for the minimum detectable effect too: a test underpowered for the effect claimed could never have found it.