Alerting

Escalation: who is woken, in what order, and who decides

Escalating for help and escalating for authority are different ladders, and most teams have only built the first one.

Teams applying this principle can also compare practical guidance on weekly timesheet template, keeping time and activity records separate from the judgement they are meant to inform.

An escalation policy is a decision about whose sleep to spend, taken in advance and in daylight, precisely so that nobody has to take it at three in the morning while also debugging something.

Most of them are built as a single ladder: primary, secondary, and then a manager. That works for one of the two reasons people escalate and not for the other.

Two ladders, not one

Escalating for help means the problem needs more skill or more hands. The person who arrives is chosen for what they know, and the trigger is that the responder is out of depth or out of time.

Escalating for authority means a decision has to be made that the responder is not entitled to make. Taking the site down deliberately. Spending money at an unusual hour. Rolling back somebody else's release. Telling customers something. Notifying a regulator, which in several jurisdictions starts a clock measured in hours.

These need different people and fire on different triggers, and a policy that only builds the first ladder produces a familiar scene: a capable engineer at four in the morning who knows exactly what should be done and does not have permission to do it.

Acknowledged is not the same as handled

Almost every system stops escalating once somebody acknowledges. The implicit claim is that acknowledgement means the problem is being worked on, and at three in the morning it frequently means a hand reached out from under a duvet.

The remedy is to keep a timer running after acknowledgement. If nothing has changed after an interval, escalate anyway. Acknowledgement should pause the ladder, not dismantle it, and treating the two as identical is how an incident quietly loses its responder for an hour.

Name roles, not people

A policy naming an individual is a policy that fails on holiday, on sick leave, and permanently on the day that person leaves. It also concentrates a system's operability in one head, which is a staffing risk disguised as a configuration.

Where a system genuinely has only one person who understands it, that is the finding, and the escalation policy is where it becomes visible. Treating it as a documentation task rather than an availability problem is the usual mistake.

primary on-callthe page arrivessomeone who can decidemoney, downtime, or a rollbacksecondaryno acknowledgement, or out of depthsomeone who can speakcustomers or a regulatorwhoever knows this systemthe failure is unfamiliarsomeone accountablethe decision outlives the incidentescalating for helpescalating for authoritydifferent people, different triggers, and teams that only build the left one
Figure 1Two ladders with different occupants. Building only the left one produces a capable engineer at four in the morning who knows what to do and is not entitled to do it.

Broadcasting is not escalating

Posting into a channel where twelve people can see it feels like raising the alarm, and it distributes responsibility to the point where nobody holds it. Each reader assumes somebody better placed is already acting, and the group's response is slower than any individual's would have been.

Escalation means naming a person or a role and handing the problem over explicitly, with an acknowledgement that the handover happened. A channel post is a useful supplement to that and a poor substitute for it.

Under-escalating is the more common failure

Teams escalate less than they should, and the reason is social rather than technical: waking a colleague feels like admitting you could not manage. So people spend an extra ninety minutes alone on something a second person would have recognised in five.

The correction is not a process change. It is a manager saying plainly, and repeatedly, that escalating early carries no penalty and that they would rather be woken. That is a statement only a manager can make, and its absence cannot be fixed with configuration.

The parts of the ladder that are outside the team

Two gaps appear reliably during real incidents. Nobody knows who holds the support contract credentials for the failing vendor, or what response time was purchased, so the first thirty minutes go on finding out. And nobody is sure who is allowed to say something publicly, so either nothing is said for too long or something is said that has to be corrected.

Both belong in the same document as the technical ladder, written before they are needed, because both are questions of authority rather than skill.

On-call is work, and in some places it is regulated

Being available outside working hours has legal status in several jurisdictions, affecting rest requirements and compensation, and the rules differ substantially between countries and sometimes by sector. This is a matter for the organisation's own advice rather than for a technical publication, and it belongs in the design of a rotation rather than being discovered afterwards.

Independently of the law, an unpaid rotation with a heavy load is a retention problem with a delay on it, and the article on alert fatigue describes the numbers that predict it.

What we cannot verify

Escalation practices vary widely by organisation size and sector, and what is described here is common practice rather than a standard; no standards body governs it. Notification obligations and working-time rules are jurisdictional, change, and nothing here is legal advice. Response times quoted in vendor support contracts are contractual claims by those vendors and are worth reading rather than assuming.

The short version

  1. An escalation policy is a decision about whose sleep to spend, taken in daylight.
  2. Escalating for help and escalating for authority need different people and triggers.
  3. Acknowledgement should pause the ladder, not dismantle it.
  4. Name roles, not individuals, or the policy fails on the first holiday.
  5. Broadcasting to a channel distributes responsibility until nobody holds it.
  6. Under-escalation is social, and only a manager can fix it.

Further context

For a primary, standards or institutional reference, see NIST guidance on continuous monitoring.

Start from the class of problem, not the list of tools

Every selection here names the situation first and the criteria second. Product names come last, and each one carries the line describing what it costs you.