Sign in Start free
ANALYTICS & REPORTING

Anomaly triage queue

Use when several metrics have alerted at once and you need to know what to look at first.

anomaly-triage-queue.md
Download .md
You are triaging a set of analytics anomalies. Treat this like an on-call queue.

Anomalies detected:
{{ANOMALY_LIST}}

Context: {{SITE_AND_BUSINESS_CONTEXT}}
Recent changes and deploys: {{RECENT_CHANGES}}

Output a table ordered by triage priority: Anomaly | Likely real or likely artefact | Revenue or decision impact if real | Fastest confirming check | Owner type (analyst, developer, paid, content) | Priority (P1 to P3).

Priority rule, apply it explicitly:
P1 = plausibly real, and if real it affects revenue or a decision this week.
P2 = plausibly real, impact is slower or smaller.
P3 = likely a measurement artefact or below normal variation.

Then:
- Group anomalies that are probably one underlying cause and name the suspected common cause.
- List anomalies that are within normal variation for this metric and should be closed with no action, with the reason.
- State what you would need to see to escalate any P3 to P1.

Rules:
- Do not raise everything to P1. If more than a third are P1, re-examine.
- If an anomaly cannot be judged without a number I did not give, say which number you need.
- No speculation dressed as diagnosis.

Fill in before running

Replace each placeholder with your own detail. The more specific you are, the less the model invents.

  • {{ANOMALY_LIST}}
  • {{SITE_AND_BUSINESS_CONTEXT}}
  • {{RECENT_CHANGES}}

Getting a better result

  1. Paste the deploy log for the same window; most simultaneous anomalies share one release.
  2. Give normal variation ranges per metric if you have them, or it will over-flag noisy ones.
  3. Close the P3s explicitly in writing - unclosed alerts are what makes people ignore alerting.

Questions about this prompt

When do I use this rather than investigating each alert in turn?

When several fired at once and they are probably one cause. Working them serially means discovering the same release four times. This groups anomalies that share an underlying cause, then orders what remains by whether it is plausibly real and whether being real changes a decision this week.

What do I need in front of me?

The anomaly list, site and business context, the deploy and change log covering the same window, and normal variation ranges per metric if you hold them. Without those ranges it over-flags the metrics that are always noisy, and the deploy log is what makes the common cause grouping work at all.

What comes back?

A table ordered by triage priority: real or artefact, impact if real, fastest confirming check, owner type and a P1 to P3 rating. The grouping and the close with no action list matter more than the ratings. One names the suspected common cause, the other clears the queue legitimately rather than by neglect.

What is the mistake that costs me here?

Leaving the P3s open. Unclosed alerts are how a team learns to ignore the channel, so write down the reason for closing each. Watch the P1 count as well: the prompt says re-examine when more than a third are P1, and a queue where everything is urgent has no priority in it.