A major financial services organisation had experienced several major incidents affecting Tier 1 services. Response was perceived as slow, and the organisation was relying on users calling in to report outages rather than finding them itself.
The tooling was not the problem. They had enterprise-grade platforms in place — Freshservice for service management, Dynatrace, SolarWinds and Azure Monitor for observability — with Azure migrations actively in progress and the expected breadth of telemetry: infrastructure, application performance, synthetic and real-user monitoring.
We were asked to establish the current maturity of the alert management capability and produce a roadmap for improvement. Three weeks, start to finish.
// THE PROBLEMTen minutes a week
A 99.9% availability target on a Tier 1 service allows around ten minutes of downtime per week. That number sets the standard for everything else: detection, escalation, resolution. It demands high levels of engineering and operational maturity, and a lot of automation.
Synthetic monitors were the only mechanism that automatically triggered a 24/7 callout. They were running every fifteen minutes.
Fifteen-minute checks cannot protect a ten-minute budget. A failure could run its entire allowance before the next check even looked — which is why every major incident that year had arrived by telephone.
"Fifteen-minute synthetic checks against a ten-minute weekly downtime budget. The arithmetic decides the outcome before anyone is paged."
// HOW WE DID ITThree weeks, four phases, six dimensions
We ran a time-boxed assessment in four short phases: mobilisation and readiness, discovery and deep dives, recommendation shaping, and final reporting — with an interim executive report at the end of the second week so leadership was not waiting for the finish.
The method was deliberately mixed. Quantitative review of system outputs and documentation, alongside daily interviews and shadow sessions with people across IT Operations, Service Management and the Platform and Delivery teams. Where evidence was requested but not available, we cross-referenced interviews rather than leaving a gap.
Capability was assessed against six dimensions of our alert management maturity model:
- Observability — detection and telemetry sources, strategy, monitoring coverage, service mapping.
- Alert quality — noise reduction, signal quality, prioritisation and business context.
- Incident response — routing, escalation, ownership and response playbooks.
- Tooling and automation — integration, automation and self-healing.
- Post incident review — alert-to-incident correlation, root cause analysis, follow-up actions.
- Continuous improvement — reporting and analytics, training, onboarding and awareness.
We set good practice at 3.5 on a scale of one to five. Most practices sat between Level 1 and Level 2, with post incident review and continuous improvement weakest and tooling and automation strongest — strongest, but underused.
// WHAT WE FOUNDEight findings, one root cause
Over 45 material observations rolled up into eight key findings. They reinforced each other, and they traced back to one thing.
Nobody knew who owned an alert. The organisation ran a modern “you build it, you run it” model, with Platform teams holding service ownership and second-line support. Yet every Platform team reported confusion over who defined and configured monitoring and alerting rules. Thresholds had been set by Service Operations from long operational experience, but without agreement from the teams who owned the services. On one occasion a Platform team’s critical synthetic alert had been disabled as part of a cost-saving exercise.
Everything else followed from that:
- Coverage gaps. Around 55% alert coverage across Tier 1 services, with almost half having no synthetic tests at all — eleven checks covering twenty Tier 1 services.
- Alerts that were not actionable. System event severities were ignored, thresholds did not map to business criticality, and alerts arrived without the context to act on them.
- Informal escalation. Major incidents were escalated in Microsoft Teams with a “scattergun” approach to finding a resolver, with runbooks thin or missing.
- Service tiers not applied. Despite the 99.9% target, 35% of Tier 1 services had no support outside core hours.
- Basic automation. Nearly every stage of incident response was manual, with no self-healing — service restarts done by hand.
- No metrics. Time to Detect and Time to Respond were not measured. Outage start was recorded as the time a user reported it, not the time the system knew.
- No central backlog. Service Operations work was spread across email, spreadsheets, Freshservice and JIRA, with roughly one man-day a week reaching alert management improvement at all.
Two numbers captured the trust problem. Of 6,390 alerts generated in six months, 1,726 led to incidents requiring action — barely a quarter. And Service Operations were performing manual “early morning checks” every day, plus watching Dynatrace dashboards, because they did not believe the alerts would tell them.
"A daily manual checklist is not a process. It is a statement of how much the team trusts its own alerting."
// WHAT WE RECOMMENDEDTwenty-nine actions, five to start now
We produced a catalogue of 29 recommendations, split into short-term (one to three months) and medium-term (three to nine months) horizons. There were no long-term items — everything on the list had merit today.
Recognising that nothing gets done if everything is attempted at once, we named the five to begin in month one.
Alongside the roadmap we ran a follow-on workshop to review the findings and help the teams shape their own improvement plan — the target being Level 3 on the maturity model, and “low service risk”.
// THE LESSONTooling is not maturity
This organisation had bought well. Dynatrace, Freshservice, SolarWinds, Azure Monitor, and the full breadth of telemetry those tools imply. The investment was sound and the foundation was genuinely strong.
What was missing was ownership — and without it, coverage drifted, thresholds went unagreed, severities were ignored, and the teams stopped trusting the output enough to rely on it. No further tooling would have fixed that.
Alert management maturity is an organisational property, not a product feature. It is also measurable, which means it can be improved deliberately rather than incident by incident.