brianletort.ai
← Library

Microsoft Azure

Verified evidenceCoordination compressedContinuous decisioning

Warden automated incident detection in IcM

A cross-service Azure incident is recognized when a human on-call, seeing only a partial service view, pieces together cascading alerts and declares the incident.

IT operations · Global Azure production (26 major services in the evaluation)

Collections: Queue eliminated · Embodied work

An editorial scene for Microsoft Azure contrasts fire alerts inside individual azure services as local symptoms appear. with continue to fire alerts; warden ingests them globally from icm. in the warden automated incident detection in icm workflow.

Executive brief

The operating-model shift, in one view.

The scarce job in a large cloud is not watching one service's alerts. It is noticing that several services are the same incident. Azure moved that recognition from on-call discussion to a global detector that pages the right people, and was faster than human declaration in about two-thirds of the incidents it caught.

AI value · Share of successfully detected incidents where Warden is faster than human declaration

Verified

Warden faster in ~68% of successfully detected incidents; median saving 21.8 NTUs, about 15% of time-to-mitigate for those cases

The public paper does not report an absolute minute delta; time is in normalized time units. The 68% figure is among successfully detected incidents, not all incidents.

Before

Fire alerts inside individual Azure services as local symptoms appear. → Triage service alerts and collaborate to recognize one cross-service incident. → Coordinates mitigation after the incident is declared.

After

Continue to fire alerts; Warden ingests them globally from IcM. → Every minute, scores recent alerts and flags potential incidents at a greater-than-90% precision operating point. → Receive a Recommended Actions notification with grouped incident-indicating alerts and an Emerging Issues dashboard.

Human boundary

Warden decides when to notify and which alerts to group. Humans retain incident declaration, severity, and mitigation. Notifications are suppressed when Warden would be slower than the human declaration already in progress.

Why it matters

A cross-service Azure incident is recognized when a human on-call, seeing only a partial service view.

How the work changed

Before

How the work ran before the change.

  1. Step 1 of 3

    Service monitors

    Fire alerts inside individual Azure services as local symptoms appear.

    ControlPer-service monitors; each team has only a partial view of the platform

  2. Step 2 of 3

    On-call engineers

    Triage service alerts and collaborate to recognize one cross-service incident.

    ControlHuman pattern recognition across teams; motivating example took nearly 50 minutes to declare

  3. Step 3 of 3

    Incident Commander

    Coordinates mitigation after the incident is declared.

    ControlIcM incident declaration

What changed

A cross-service Azure incident is recognized when a human on-call, seeing only a partial service view.

Decision rightSelection moves from a fixed rule to the model

After

How the same work runs now.

  1. Step 1 of 3

    Service monitors

    Continue to fire alerts; Warden ingests them globally from IcM.

    ControlSelected monitors with high weighted mutual information versus historical incidents

  2. Step 2 of 3

    Warden detector

    Every minute, scores recent alerts and flags potential incidents at a greater-than-90% precision operating point.

    ControlPrecision held above 90% (recall ~58%, F1 0.71); late detections relative to humans are not notified

  3. Step 3 of 3

    On-call engineers

    Receive a Recommended Actions notification with grouped incident-indicating alerts and an Emerging Issues dashboard.

    ControlHumans still declare and mitigate; Warden does not auto-close or auto-mitigate the incident

Process model built from the published workflow evidence for Microsoft Azure. Every step, actor, and control appears in full below.
Every step, actor, and control

Exception path

False positives are constrained by a >90% precision threshold. Incidents Warden misses wait for human declaration. Not all incidents are covered by monitors; the authors flag uncovered incidents as future work.

Work removed

  • Ad-hoc multi-team discussion as the primary way to recognize a cross-service incident
  • Each on-call independently deciding whether a local alert is isolated or part of a larger event

Decision authority

Warden decides when to notify and which alerts to group. Humans retain incident declaration, severity, and mitigation. Notifications are suppressed when Warden would be slower than the human declaration already in progress.

Before

  1. 01

    Service monitors

    Fire alerts inside individual Azure services as local symptoms appear.

    Control: Per-service monitors; each team has only a partial view of the platform

  2. 02

    On-call engineers

    Triage service alerts and collaborate to recognize one cross-service incident.

    Control: Human pattern recognition across teams; motivating example took nearly 50 minutes to declare

  3. 03

    Incident Commander

    Coordinates mitigation after the incident is declared.

    Control: IcM incident declaration

After

  1. 01

    Service monitors

    Continue to fire alerts; Warden ingests them globally from IcM.

    Control: Selected monitors with high weighted mutual information versus historical incidents

  2. 02

    Warden detector

    Every minute, scores recent alerts and flags potential incidents at a greater-than-90% precision operating point.

    Control: Precision held above 90% (recall ~58%, F1 0.71); late detections relative to humans are not notified

  3. 03

    On-call engineers

    Receive a Recommended Actions notification with grouped incident-indicating alerts and an Emerging Issues dashboard, then prioritize and start cross-team collaboration.

    Control: Humans still declare and mitigate; Warden does not auto-close or auto-mitigate the incident

Work that left the path

  • Ad-hoc multi-team discussion as the primary way to recognize a cross-service incident
  • Each on-call independently deciding whether a local alert is isolated or part of a larger event

Human role before

Detect that locally visible alerts are one cross-service incident, then engage peers and an incident commander.

Human role after

Triage Warden-prioritized emerging issues, confirm the grouping, and run mitigation. Detection of the cross-service pattern is no longer the on-call's first job.

AI roleFrom a global alert stream, detect a potential incident, extract incident-indicating alert groups, and notify the relevant on-calls so they can collaborate sooner.

Outcomes

Share of successfully detected incidents where Warden is faster than human declaration

Verified

On-call incident declaration time for the same incidentsWarden faster in ~68% of successfully detected incidents; median saving 21.8 NTUs, about 15% of time-to-mitigate for those cases

18 months of Azure IcM data starting October 2018 (last 2 months held out for test); production deployment about 3 months through May 2020 · 26 major Azure services accounting for ~72% of Azure incidents; >10 million alerts; ~240 GB

The public paper does not report an absolute minute delta; time is in normalized time units. The 68% figure is among successfully detected incidents, not all incidents. Comparison is detection latency versus historical human declaration for the same incident, not operating cost or full mitigation time. No public 2026 confirmation that Warden remains the production detector under that name.

What leaders can reuse

Anti-pattern

Reporting '68% faster than human' as a minute-scale MTTR cut, or implying Warden closed incidents without an on-call.

Questions

  1. 01Is the outcome time-to-detect versus the same incident's human declaration, or a modeled counterfactual?
  2. 02What precision floor is high enough that on-calls will trust the page?
  3. 03If the detector is late, do we suppress the notification, or create a second incident?

Portability conditions

  • A shared incident platform that already stores alerts across services
  • Willingness to operate at high precision and accept missed detections
  • On-call UX that shows grouped alerts, not a new ticket pile
  • A human declaration path that remains authoritative when the model is late or wrong

Reputation risk

low

Evidence and authority

What the public record supports.

Current · updated

1 peer reviewed, 1 primary; publication outcomes are verified.

Bundle 1.0.0 · reviewed 2026-08-23 · stable ID 7405cae4f09ed2d4

Related transformations

More in IT operations

Sources

Read the evidence, freshness, caveat, and version policy.