It is 3 a.m. The on-call phone buzzes for the eleventh time tonight. Disk at 81%. CPU spike on a batch node. A health check flapped and recovered. None of it needs action. When a real outage alert arrives at 4 a.m., it gets the same tired glance as the others.
That is alert fatigue, and it is one of the most common reasons teams come to us. This playbook explains why it happens and how to fix it.
What is alert fatigue?
Alert fatigue is the desensitisation that happens when people receive too many alerts, especially false positives or non-actionable ones. Responders begin to ignore, mute or delay responding to alerts, which increases the risk of missing a real incident. It also contributes directly to on-call burnout and attrition.
What causes alert fatigue?
Alert inhibition and correlation example
Illustration in progress
- Alerting on causes, not symptoms. CPU, memory and disk alerts fire constantly, but users may never notice.
- Static thresholds that ignore normal daily and weekly patterns.
- Duplicate alerts from multiple tools watching the same thing.
- Cascading alerts where one failure triggers dozens of downstream alerts.
- No ownership. Alerts go to a shared channel where everyone assumes someone else will look.
- Alerts nobody ever reviewed or deleted. Alert rules accumulate over years.
How do you measure alert fatigue?
Before you fix it, measure it for two to four weeks:
Alert noise reduced by 70 to 80 percent
Illustration in progress
| Metric | What it tells you |
|---|---|
| Alerts per on-call shift | Overall volume |
| Percentage of alerts acted on | Signal-to-noise ratio |
| Top 10 noisiest alert rules | Where to start |
| Alerts auto-resolved within 5 minutes | Flapping candidates |
| Pages outside working hours | Burnout risk |
If fewer than half of your alerts lead to action, you have a noise problem.
If fewer than half of your alerts lead to action, you have a noise problem.
A step-by-step playbook to reduce alert fatigue
1. Define what deserves a page
A page should be urgent, actionable and need a human. Everything else becomes a ticket, a dashboard or a daily report.
2. Alert on symptoms tied to SLOs
Replace dozens of infrastructure alerts with a few alerts on what users experience: error rate, latency and availability. Use SLO burn-rate alerts that fire when you are consuming your error budget too fast. See our SLO and error budget guide.
3. Delete or downgrade the noisiest rules
Take the top 10 noisiest alerts. For each, ask: did anyone take action in the last month? If not, delete it or downgrade it to a non-paging notification.
4. Correlate, group and inhibit
Use grouping to combine related alerts into one notification. Use inhibition rules so that when a parent component is down, alerts from dependent services are suppressed.
5. Use dynamic baselines where patterns vary
For metrics with daily or seasonal patterns, anomaly detection can replace static thresholds that fire every Monday morning.
6. Route to owners and attach runbooks
Every alert should name an owning team and link to a runbook explaining what to check and do first.
7. Review alerts every sprint
Add a 15-minute alert review to the on-call handover. Tune, delete or improve anything that fired without action.
What does a good alert look like?
A good alert includes: a clear title describing user impact, severity, the affected service and owner, a link to the relevant dashboard and traces, and a runbook link. "Checkout error rate 4.2% (SLO 99.5%) – burning error budget 10x – payments team – runbook" is far more useful than "HTTP 5xx high on node-17".
How Crozaint approaches alert fatigue
High-signal alerting is a core deliverable of Crozaint's 8-week observability engagement. In the Operationalize phase (weeks 5–6), we design SLOs, then engineer alert rules with correlation and inhibition that typically reduce alert noise by 70–80%.
We build on Grafana Cloud with OpenTelemetry instrumentation, and use Grafana's machine learning features for anomaly detection and Sift to speed up investigation once an alert fires. AI assists, but humans stay in control: nothing acts autonomously without approval.
Common mistakes to avoid
- Lowering thresholds after every incident, adding noise
- Paging on warnings that need no immediate action
- Sending all alerts to one shared channel
- Deleting alerts without checking SLO coverage first
- Treating alert tuning as a one-time project
Conclusion
Alert fatigue is a design problem, not a people problem. Page only on urgent, actionable symptoms, correlate the rest, assign owners and review every sprint.
On-call team drowning in alerts? Book a 30-minute discovery call and let us review your alerting with you.
