Escalation
Escalation policies define what happens after an alert is created—who is notified first, when to re‑notify, and when to involve additional teams or management.
Use this guide to learn how escalation policies route alerts to the appropriate responders and help ensure that incidents are acknowledged within the expected timeframe.
Escalation Patterns
Common patterns include:
- Single‑team on‑call – alerts go to a primary on‑call engineer, then to a backup if unacknowledged.
- NOC‑first – a network operations center triages alerts and forwards incidents to application teams when needed.
- Functional escalation – infrastructure teams handle platform issues; application teams handle service‑specific issues; security teams handle security‑related alerts.
Timers and Re‑Notification
Escalation policies typically define:
- How long to wait for acknowledgment before escalating.
- Whether to re‑notify the same person or move to the next contact.
- When to auto‑resolve alerts that have been stable for some time.
Use severity to influence these timers — critical alerts may escalate more quickly than warning alerts.
Integration with On‑Call and ITSM
Escalation often involves:
- On‑Call tools such as PagerDuty or equivalent systems.
- Ticketing systems (for example, ServiceNow or Jira) where incidents are tracked.
Use integrations documented under Other Integrations to connect Unified Monitoring alerts to your existing escalation workflows.
Best Practices
- Maintain clear ownership for each service (for example, an
Application Ownerfield in CMDB) and use it in routing and escalation. - Keep escalation trees as simple as possible and align them with your on‑call rotation structures.
- Test escalation flows periodically (for example, via synthetic alerts) to ensure they still work.