back_to_insights

// CASE STUDY · SERVICE OPERATIONS TRANSFORMATION

An alert management maturity review for a major Financial Services organisation

Despite a 99.9% availability target, every major incident that year was first reported directly by customers. Axiologik were engaged to understand why.

IMAGE: PHOTO BY ZULFUGAR KARIMOV ON UNSPLASH

A major financial services organisation had experienced several major incidents affecting Tier 1 services. Response was perceived as slow, and the organisation was relying on users calling in to report outages rather than finding them itself.

The tooling was not the problem. They had enterprise-grade platforms in place — Freshservice for service management, Dynatrace, SolarWinds and Azure Monitor for observability — with Azure migrations actively in progress and the expected breadth of telemetry: infrastructure, application performance, synthetic and real-user monitoring.

We were asked to establish the current maturity of the alert management capability and produce a roadmap for improvement. Three weeks, start to finish.

// THE PROBLEMTen minutes a week

A 99.9% availability target on a Tier 1 service allows around ten minutes of downtime per week. That number sets the standard for everything else: detection, escalation, resolution. It demands high levels of engineering and operational maturity, and a lot of automation.

Synthetic monitors were the only mechanism that automatically triggered a 24/7 callout. They were running every fifteen minutes.

Fifteen-minute checks cannot protect a ten-minute budget. A failure could run its entire allowance before the next check even looked — which is why every major incident that year had arrived by telephone.

"Fifteen-minute synthetic checks against a ten-minute weekly downtime budget. The arithmetic decides the outcome before anyone is paged."

// HOW WE DID ITThree weeks, four phases, six dimensions

We ran a time-boxed assessment in four short phases: mobilisation and readiness, discovery and deep dives, recommendation shaping, and final reporting — with an interim executive report at the end of the second week so leadership was not waiting for the finish.

The method was deliberately mixed. Quantitative review of system outputs and documentation, alongside daily interviews and shadow sessions with people across IT Operations, Service Management and the Platform and Delivery teams. Where evidence was requested but not available, we cross-referenced interviews rather than leaving a gap.

Capability was assessed against six dimensions of our alert management maturity model:

  • Observability — detection and telemetry sources, strategy, monitoring coverage, service mapping.
  • Alert quality — noise reduction, signal quality, prioritisation and business context.
  • Incident response — routing, escalation, ownership and response playbooks.
  • Tooling and automation — integration, automation and self-healing.
  • Post incident review — alert-to-incident correlation, root cause analysis, follow-up actions.
  • Continuous improvement — reporting and analytics, training, onboarding and awareness.

We set good practice at 3.5 on a scale of one to five. Most practices sat between Level 1 and Level 2, with post incident review and continuous improvement weakest and tooling and automation strongest — strongest, but underused.

// WHAT WE FOUNDEight findings, one root cause

Over 45 material observations rolled up into eight key findings. They reinforced each other, and they traced back to one thing.

Nobody knew who owned an alert. The organisation ran a modern “you build it, you run it” model, with Platform teams holding service ownership and second-line support. Yet every Platform team reported confusion over who defined and configured monitoring and alerting rules. Thresholds had been set by Service Operations from long operational experience, but without agreement from the teams who owned the services. On one occasion a Platform team’s critical synthetic alert had been disabled as part of a cost-saving exercise.

Everything else followed from that:

  • Coverage gaps. Around 55% alert coverage across Tier 1 services, with almost half having no synthetic tests at all — eleven checks covering twenty Tier 1 services.
  • Alerts that were not actionable. System event severities were ignored, thresholds did not map to business criticality, and alerts arrived without the context to act on them.
  • Informal escalation. Major incidents were escalated in Microsoft Teams with a “scattergun” approach to finding a resolver, with runbooks thin or missing.
  • Service tiers not applied. Despite the 99.9% target, 35% of Tier 1 services had no support outside core hours.
  • Basic automation. Nearly every stage of incident response was manual, with no self-healing — service restarts done by hand.
  • No metrics. Time to Detect and Time to Respond were not measured. Outage start was recorded as the time a user reported it, not the time the system knew.
  • No central backlog. Service Operations work was spread across email, spreadsheets, Freshservice and JIRA, with roughly one man-day a week reaching alert management improvement at all.

Two numbers captured the trust problem. Of 6,390 alerts generated in six months, 1,726 led to incidents requiring action — barely a quarter. And Service Operations were performing manual “early morning checks” every day, plus watching Dynatrace dashboards, because they did not believe the alerts would tell them.

"A daily manual checklist is not a process. It is a statement of how much the team trusts its own alerting."

// WHAT WE RECOMMENDEDTwenty-nine actions, five to start now

We produced a catalogue of 29 recommendations, split into short-term (one to three months) and medium-term (three to nine months) horizons. There were no long-term items — everything on the list had merit today.

Recognising that nothing gets done if everything is attempted at once, we named the five to begin in month one.

Publish roles and responsibilities
Who owns definition, coverage, thresholds and configuration — enforced through Service Acceptance, with the engagement model published alongside. Every other improvement depends on this.
Increase synthetic frequency to 2–5 minutes
The only mechanism that triggers 24/7 callout must fire before the SLA is spent. There is a cost; against a 99.9% target it is justifiable.
Implement core alert management metrics
Time to Detect and Time to Respond, with KPIs and trends — measured from system data, not from when a customer called.
Establish timely escalation workflows
Named SREs and engineers available for third-line support on every Tier 1 service, with on-call schedules where none exist.
Put all Service Operations work on one backlog
Project tasks, Platform team requests and post-incident actions in one place, prioritised against each other and categorised for capacity modelling.

Alongside the roadmap we ran a follow-on workshop to review the findings and help the teams shape their own improvement plan — the target being Level 3 on the maturity model, and “low service risk”.

// THE LESSONTooling is not maturity

This organisation had bought well. Dynatrace, Freshservice, SolarWinds, Azure Monitor, and the full breadth of telemetry those tools imply. The investment was sound and the foundation was genuinely strong.

What was missing was ownership — and without it, coverage drifted, thresholds went unagreed, severities were ignored, and the teams stopped trusting the output enough to rely on it. No further tooling would have fixed that.

Alert management maturity is an organisational property, not a product feature. It is also measurable, which means it can be improved deliberately rather than incident by incident.

// PRODUCT · DIGITAL SERVICE OPERATIONS HEALTHCHECK

Find out where your operations actually stand.

The Digital Service Operations Healthcheck is a structured, expert-led review of how effectively your services are run — the operating model, practices, tooling and service performance outcomes.

You get an honest picture of where you stand, where the problems are and what to address in what order — all underpinned by actual service performance data.

WHAT THE HEALTHCHECK LOOKS AT

01
Operating model and leadership

How operations is structured, who is accountable, and where product and operations meet.

02
Service performance

What is measured, whether it reflects the business, and what the numbers are hiding.

03
Stability and recovery

Incident and problem practice, root cause discipline, and how services behave under strain.

04
Change and release control

How much governance is manual, how much is automated, and what change is costing you.

05
Observability and SRE maturity

Instrumentation, service level objectives and whether reliability is engineered or watched for.

06
AI readiness and AI in production

Where AI would help operations — and how the AI services you already run are governed and supported.