Back to search
Yochana Linkedin · Posted 2d ago

Major Incident Manager

Wilmington

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Job Title: Project Manager (Major Incident Management & NOC )

Location: Onsite – Wilmington, DE (Day1 Onsite)

Employment Type: Full-time & Contract

Experience: 10+ years in IT Operations / NOC / Major Incident Management



Role Summary:

The Project Manager is responsible for Major Incident Management & NOC Teams .

This role leads the team for MIM and NOC functions who drive Major Incident (P1/P2) execution, ensures rapid service restoration, and continuously improves operational maturity through problem management, automation, observability enhancements, and SLA governance.

The role requires a mix of strong incident leadership, technical depth across infrastructure and applications, and people/process management to ensure stability, availability, and performance across critical services.

Key Responsibilities:

A) Manage Team of Major Incident Managers (Command & Control)

Own the Major Incident (P1/P2) process from detection to resolution, including war-room leadership, stakeholder updates, and closure.

Ensure structured triage, containment, workaround, and restoration.

Drive cross-functional coordination (App, Infra, Network, Security, DB, Cloud, Vendor teams) to reduce MTTR.

Ensure high-quality incident communications: executive summaries, impact analysis, ETAs, customer/business comms.

Lead and facilitate Post Incident Reviews (PIR/RCA); ensure actionable corrective/preventive actions (CAPA).

Identify recurring issues and trigger Problem Management with measurable reduction plans.

ITIL v4 Foundation (preferred).

Reduced MTTD and MTTR for P1/P2 incidents.

Improved SLA compliance and reduction in escalation breaches.

Reduced repeat incidents via problem management and preventive actions.

Improved alert quality: lower false positives, better signal-to-noise ratio.

Strong PIR/RCA compliance: on-time RCAs with measurable preventive outcomes.

Improved NOC operational maturity: SOP adherence, shift handover quality, audit readiness.

B) NOC Leadership & Operations

Manage the NOC team responsible for 24x7 monitoring, alert triage, event correlation, escalation, and ticket quality.

Establish/maintain standard operating procedures (SOPs), runbooks, escalation matrices, and on-call models.

Ensure NOC meets SLAs/OLAs, improves alert fidelity, and reduces noise through tuning and automation.

Manage handover governance between shifts; maintain service continuity and operational hygiene.

C) Service Reliability & Continuous Improvement

D) Stakeholder & Vendor Management

E) Managerial / Leadership Skills (Must Have)


Technical Skills (Must Have):

A) Monitoring / Observability

B) Incident / ITSM Platforms (Good to Have )

Experience designing workflows, SLAs/OLAs, routing rules, and automation integrations.

Hands-on experience with NOC tooling and observability platforms such as:Splunk / ELK, Datadog, Dynatrace, New Relic, AppDynamics,Prometheus/Grafana, CloudWatch/Azure Monitor

C) Infrastructure & Platform Breadth

Solid understanding across:

Windows/Linux administration basics

Network fundamentals (DNS, DHCP, TCP/IP, routing, load balancers, firewalls)

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search