Sign in Start free
STRATEGY

Message testing plan

Use when two or more messaging directions are on the table and the argument cannot be settled by opinion.

message-testing-plan.md
Download .md
You are a research lead designing the cheapest test that would actually change our minds.

Competing messages: {{MESSAGE_OPTIONS}}
Audience: {{AUDIENCE}}
Traffic or list size available: {{AUDIENCE_VOLUME}}
Time and budget: {{TIME_AND_BUDGET}}
What we would do differently depending on the result: {{DECISION_AT_STAKE}}

Produce:
1. State the decision the test informs. If the answer would not change what we do, say the test is not worth running and stop.
2. For each message, the underlying belief it depends on. That belief, not the wording, is what we are testing.
3. A table with columns: Test method | What it measures | Cost | Time to result | Sample needed | What it cannot tell us.
4. Recommend one method given the volume and budget, and say why the cheaper options are insufficient here.
5. The exact success criterion: the metric, the threshold, and the minimum sample. If the available volume cannot reach that sample, say so plainly and propose a qualitative alternative.
6. Three ways this test could mislead us, and the guard for each.

Constraints: do not propose a split test where the traffic cannot reach significance. Do not quote a required sample size as a precise figure without showing the assumed baseline rate and effect size. No em dashes.

Fill in before running

Replace each placeholder with your own detail. The more specific you are, the less the model invents.

  • {{MESSAGE_OPTIONS}}
  • {{AUDIENCE}}
  • {{AUDIENCE_VOLUME}}
  • {{TIME_AND_BUDGET}}
  • {{DECISION_AT_STAKE}}

Getting a better result

  1. Step 1 kills about half of proposed tests, which saves more time than running them well.
  2. Test the underlying belief rather than the headline wording; wording tests rarely move anything.
  3. If volume is low, five customer calls beat an underpowered split test - let it recommend that.

Questions about this prompt

When should I use this rather than just running the A/B test?

Before you run one. Step 1 asks what decision the result changes and stops if the answer is nothing, which kills a large share of proposed tests. It also checks whether your traffic can reach the sample the test would need, before you spend three weeks proving nothing either way.

What do I need in order to design the test?

The competing messages, the audience, an honest volume figure, time and budget, and above all what you would do differently depending on the result. Without that last input step 1 cannot do its job. Volume means the traffic that will actually see the test, not total monthly sessions.

What comes back, and what am I really testing?

The decision the test informs, the belief sitting under each message, a method table including what each method cannot tell you, one recommendation with the case against the cheaper options, an exact success criterion with threshold and minimum sample, and three ways it could mislead you. The belief, not the wording, is what you test.

What is the mistake that wastes three weeks?

Running an underpowered split test because it feels more rigorous than talking to people. Where volume cannot reach the sample, the prompt says so and offers a qualitative alternative. Note that any required sample figure depends on the assumed baseline rate and effect size, which is why it shows those rather than a bare number.