D

Major Incident Manager

DXC Technology

Riyadh, Riyadh Province, Saudi Arabia · Full Time

Be the first to apply

Experience
Any
Salary
Openings
1
Posted
منذ ساعتين
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Role Overview

As a Major Incident Manager, you will lead 24/7 incident management efforts to minimize disruptions and business impact, ensuring swift resolution and communication during critical service incidents.

Key Responsibilities

  • Conduct round-the-clock incident management to reduce service interruptions and mitigate business effects.
  • Log and oversee incidents caused by service degradation, outages, or operational alerts.
  • Set up and handle communication channels and Incident Bridges for critical issues.
  • Coordinate technical troubleshooting, restoration activities, and stakeholder engagement efficiently.
  • Facilitate quick decisions and clear operational blockages during incident handling.
  • Lead efforts to restore services promptly while following Incident Management procedures.
  • Organize and moderate Major Incident Review (MIR) meetings post-incident resolution.
  • Record incident timelines, triage steps, recovery measures, lessons learned, and improvement possibilities.
  • Manage stakeholder involvement in review meetings and finalize MIR reports.
  • Oversee Root Cause Analysis (RCA) for major and recurring incidents involving internal teams, OEMs, third parties, and stakeholders.
  • Identify proactively and create problem records based on incident trends and risks.
  • Perform in-depth problem analysis and impact evaluations.
  • Create and handle problem records documenting root causes, workarounds, corrective actions, and permanent fixes.
  • Monitor ongoing problems and track resolution progress within agreed service levels.
  • Lead investigative efforts to uncover underlying incident causes.
  • Collaborate with technical teams, vendors, OEMs, and regulatory stakeholders to implement lasting solutions.
  • Champion proactive remediation initiatives to prevent recurring incidents.
  • Ensure implementation and validation of corrective and preventive measures.
  • Coordinate temporary fixes to maintain service continuity during permanent solution development.
  • Validate the success of workarounds and permanent resolutions.
  • Generate regular KPI reports including problem trends, RCA status, resolution timelines, recurring issues, incident counts, and SLA compliance.
  • Ensure adherence to Incident and Problem Management policies, operational procedures, and regulatory standards.
  • Identify opportunities for process enhancements and automation, providing recommendations to the CSI team.
  • Drive initiatives focused on minimizing incident recurrence and boosting service availability.
  • Update knowledge bases with validated root causes, known errors, workarounds, and permanent solutions.
  • Ensure documentation and communication of lessons learned from major incidents across support teams.

Required Technical Skills

  • Expertise in Major Incident Management
  • Proficiency in Root Cause Analysis (RCA) techniques
  • Strong understanding of Incident and Problem Management practices
  • Knowledge of ITIL Service Management Framework
  • Ability to analyze trends, report metrics, and perform operational analytics
  • Understanding of infrastructure, applications, cloud services, networks, databases, and security operations
  • Familiarity with SLAs, KPIs, and service performance management

Performance Indicators

  • Decrease in recurring incidents
  • High percentage of major incidents with complete RCA
  • Timely closure of problems and major incidents within SLA targets
  • Reduced Mean Time to Identify Root Cause (MTTRCA)
  • Number of proactive problems identified and resolved
  • Reduction in incident volume through permanent resolutions
🤖
Online · instant AI help