Anomaly triage queue
Use when several metrics have alerted at once and you need to know what to look at first.
Fill in before running
Replace each placeholder with your own detail. The more specific you are, the less the model invents.
- {{ANOMALY_LIST}}
- {{SITE_AND_BUSINESS_CONTEXT}}
- {{RECENT_CHANGES}}
Getting a better result
- Paste the deploy log for the same window; most simultaneous anomalies share one release.
- Give normal variation ranges per metric if you have them, or it will over-flag noisy ones.
- Close the P3s explicitly in writing - unclosed alerts are what makes people ignore alerting.
Questions about this prompt
When do I use this rather than investigating each alert in turn?
When several fired at once and they are probably one cause. Working them serially means discovering the same release four times. This groups anomalies that share an underlying cause, then orders what remains by whether it is plausibly real and whether being real changes a decision this week.
What do I need in front of me?
The anomaly list, site and business context, the deploy and change log covering the same window, and normal variation ranges per metric if you hold them. Without those ranges it over-flags the metrics that are always noisy, and the deploy log is what makes the common cause grouping work at all.
What comes back?
A table ordered by triage priority: real or artefact, impact if real, fastest confirming check, owner type and a P1 to P3 rating. The grouping and the close with no action list matter more than the ratings. One names the suspected common cause, the other clears the queue legitimately rather than by neglect.
What is the mistake that costs me here?
Leaving the P3s open. Unclosed alerts are how a team learns to ignore the channel, so write down the reason for closing each. Watch the P1 count as well: the prompt says re-examine when more than a third are P1, and a queue where everything is urgent has no priority in it.