Major Incident Manager
Indexed description
Job Title: Project Manager (Major Incident Management & NOC )
Location: Onsite – Wilmington, DE (Day1 Onsite)
Employment Type: Full-time & Contract
Experience: 10+ years in IT Operations / NOC / Major Incident Management
Role Summary:
The Project Manager is responsible for Major Incident Management & NOC Teams .
This role leads the team for MIM and NOC functions who drive Major Incident (P1/P2) execution, ensures rapid service restoration, and continuously improves operational maturity through problem management, automation, observability enhancements, and SLA governance.
The role requires a mix of strong incident leadership, technical depth across infrastructure and applications, and people/process management to ensure stability, availability, and performance across critical services.
Key Responsibilities:
A) Manage Team of Major Incident Managers (Command & Control)
Own the Major Incident (P1/P2) process from detection to resolution, including war-room leadership, stakeholder updates, and closure.
Ensure structured triage, containment, workaround, and restoration.
Drive cross-functional coordination (App, Infra, Network, Security, DB, Cloud, Vendor teams) to reduce MTTR.
Ensure high-quality incident communications: executive summaries, impact analysis, ETAs, customer/business comms.
Lead and facilitate Post Incident Reviews (PIR/RCA); ensure actionable corrective/preventive actions (CAPA).
Identify recurring issues and trigger Problem Management with measurable reduction plans.
ITIL v4 Foundation (preferred).
Reduced MTTD and MTTR for P1/P2 incidents.
Improved SLA compliance and reduction in escalation breaches.
Reduced repeat incidents via problem management and preventive actions.
Improved alert quality: lower false positives, better signal-to-noise ratio.
Strong PIR/RCA compliance: on-time RCAs with measurable preventive outcomes.
Improved NOC operational maturity: SOP adherence, shift handover quality, audit readiness.
B) NOC Leadership & Operations
Manage the NOC team responsible for 24x7 monitoring, alert triage, event correlation, escalation, and ticket quality.
Establish/maintain standard operating procedures (SOPs), runbooks, escalation matrices, and on-call models.
Ensure NOC meets SLAs/OLAs, improves alert fidelity, and reduces noise through tuning and automation.
Manage handover governance between shifts; maintain service continuity and operational hygiene.
C) Service Reliability & Continuous Improvement
D) Stakeholder & Vendor Management
E) Managerial / Leadership Skills (Must Have)
Technical Skills (Must Have):
A) Monitoring / Observability
B) Incident / ITSM Platforms (Good to Have )
Experience designing workflows, SLAs/OLAs, routing rules, and automation integrations.
Hands-on experience with NOC tooling and observability platforms such as:Splunk / ELK, Datadog, Dynatrace, New Relic, AppDynamics,Prometheus/Grafana, CloudWatch/Azure Monitor
C) Infrastructure & Platform Breadth
Solid understanding across:
Windows/Linux administration basics
Network fundamentals (DNS, DHCP, TCP/IP, routing, load balancers, firewalls)
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search