Alerting

Alert fatigue is a measurable failure, not a mood

A predictable response to a signal that is usually wrong, studied where the stakes are lives, and computable from records you already hold.

Teams applying this principle can also compare practical guidance on employee time tracking app, keeping time and activity records separate from the judgement they are meant to inform.

Alert fatigue is discussed as though it were a mood, or a sign that a team lacks discipline. It is neither. It is a predictable behavioural response to a signal with low precision, it has been studied seriously in a field where the consequences are measured in lives, and it can be quantified from data most teams already have.

Where the evidence comes from

Hospitals ran this experiment at scale without meaning to. Monitoring devices at the bedside produce enormous numbers of alarms, the great majority of which do not require intervention, and staff responded the way anyone responds to a signal that is usually wrong: more slowly, and sometimes not at all.

This was taken seriously enough that accreditation bodies issued formal safety warnings and made alarm management a stated priority, on the basis of harm that had already occurred. The source is a regulator addressing patient safety rather than anybody selling monitoring software, which is why it is worth naming rather than gesturing at.

The mechanism transfers exactly. A pager that is usually wrong trains its recipient to treat pages as probably wrong, and that training does not distinguish between the alert that was noise and the one that was not.

The arithmetic behind it

The article on thresholds sets out why: when the event is rare, even an accurate detector produces mostly false alarms. Precision, the share of alerts that were real, is what determines the learned response, and precision is dominated by the base rate rather than by the quality of the detector.

This is why "make the detector better" is such a weak lever on its own and why reducing the number of things that alert at all is such a strong one.

What to measure, since it is measurable

Five numbers, all computable from existing records, none requiring a new product.

Pages per person per week, and separately, pages outside working hours per person per month. Volume is the crude measure and the one that predicts attrition.

The share of alerts that resulted in any action. This is the precision figure and the one that matters. A system where four alerts in five are closed without anyone doing anything is not a system with a tired team.

Time to acknowledge, tracked over months rather than as an average. Fatigue shows up as a slope: the same alerts being answered a little later each month.

The share that resolved themselves before anyone touched them, which counts alerts that never needed a human at all.

And volume by alert name. This one usually ends the analysis, because a handful of rules almost always account for most of the traffic.

One month of pagesled to an actionclosed without onetime to acknowledgeweeks
Figure 1A month of pages, of which a small proportion led to any action, and acknowledgement times lengthening across the same period. Both figures are computable from records most teams already keep.

The concentration is the good news

In practice the distribution is extremely uneven. Five rules generate half the pages, and they are generally not the important five. That means the problem is tractable in an afternoon rather than requiring a programme: list the top offenders, and for each one decide whether to fix the underlying cause, add a duration requirement, group it with its siblings, route it somewhere that is not a person, or delete it.

Most teams have never looked at that list, which is the only reason the situation persists.

Set a budget, and treat breaching it as work

A stated maximum, for example a number of pages per on-call shift beyond which the rotation is considered to be failing, converts a grumble into a tracked condition. The important part is what happens when it is breached: the response has to be work to reduce the volume, not tolerance of it, and not quiet adjustment of the thresholds, which is the ratchet described in the previous article.

The cost is larger than the interruption

An interruption during the day costs the task that was in progress and the time to rebuild the context, which is considerably longer than the alert took to read. At night it costs the remainder of the sleep and a degraded day after.

Across months it costs the willingness to carry the pager, which is how a monitoring problem becomes a staffing problem. That connection is rarely made in the meeting where an extra alert is added, because the alert has a visible benefit and a cost paid by somebody else later.

One caution about measuring it

The precision figure depends on people recording what they did with an alert. The moment that record is used to judge the individuals rather than the alerting system, the data stops being about alerts, for exactly the reason set out in the article on measures becoming targets.

Keep the measurement pointed at the rules, not at the responders.

What we cannot verify

The clinical evidence concerns a different environment with different stakes, and the transfer of the mechanism is an argument rather than a measurement. Figures on interruption cost circulate widely in this industry with weak attribution and we reproduce none. Product claims to reduce alert noise come from their vendors. What your own precision is can be computed from your own incident records this week, and almost nobody has.

The short version

  1. Fatigue is a predictable response to a low-precision signal, not a lack of discipline.
  2. Clinical alarm safety is the strongest evidence, and it comes from regulators, not vendors.
  3. Precision is dominated by the base rate, so reducing what alerts beats improving detectors.
  4. Five computable numbers, starting with the share of alerts that led to any action.
  5. A handful of rules generate most of the volume, which makes it an afternoon's work.
  6. Point the measurement at the rules; aimed at responders, the data becomes fiction.

Further context

For a primary, standards or institutional reference, see the CNCF project catalogue.

Start from the class of problem, not the list of tools

Every selection here names the situation first and the criteria second. Product names come last, and each one carries the line describing what it costs you.