NOC Technical Lead (L3)
Indexed description
Roles & Responsibilities
- Provide technical leadership for the 24×7 operation of COLO, ISP, AI GPU Cloud, and Data Center infrastructure.
- Lead the resolution of critical (P1/P2) incidents, major service outages, and complex technical escalations.
- Perform advanced troubleshooting and Root Cause Analysis (RCA) across network, compute, storage, cloud, virtualization, and security infrastructure.
- Own Problem Management activities by identifying recurring issues and implementing permanent corrective and preventive actions.
- Design, optimize, and continuously improve Data Center, cloud, and monitoring architectures to enhance service availability and operational efficiency.
- Develop and implement automation, orchestration, and operational tooling to reduce manual effort and improve service reliability.
- Define and maintain operational standards, monitoring strategies, alert thresholds, SOPs, runbooks, and knowledge base documentation.
- Review and approve technical procedures, operational documentation, and standard operating practices & Support Compliance Audits.
- Drive capacity planning, performance optimization, scalability, and high-availability initiatives for Data Center and AI GPU infrastructure.
- Provide technical governance for enterprise networking, cloud connectivity, security infrastructure, and observability platforms.
- Collaborate with Infrastructure, Cloud, Security, Platform Engineering, Application, Facilities, and Vendor teams to resolve complex technical issues.
- Lead technical support during planned maintenance, software upgrades, migrations, disaster recovery, and business continuity exercises.
- Participate in architecture reviews, technology evaluations, and infrastructure lifecycle planning.
- Ensure compliance with security policies, governance standards, operational procedures, and SLA/OLA commitments.
- Lead operational readiness reviews for new services, infrastructure deployments, and production transitions.
- Mentor and provide technical guidance to NOC L1/L2 engineers through knowledge transfer, technical reviews, and operational coaching.
- Provide expert support for AI GPU clusters, InfiniBand/Ethernet fabrics, Kubernetes platforms, distributed storage, and high-performance networking.
- Optimize monitoring, alert correlation, and observability across GPU, compute, storage, and network infrastructure using enterprise monitoring platforms.
- Lead service restoration during large-scale infrastructure failures and coordinate technical bridge calls with OEMs, cloud providers, and engineering teams.
- Analyze infrastructure performance trends and recommend architecture enhancements to improve scalability, resilience, and operational efficiency.
Relevant Experience
- 8–15 years of experience in Enterprise Network Operations, Data Center Operations, Cloud Infrastructure, or Managed Services environments.
- Expert knowledge of Data Center networking, including LAN/WAN, TCP/IP, MPLS, EVPN-VXLAN, Spine-Leaf architectures, BGP, OSPF, IS-IS, VLANs, STP/MSTP, VRRP/HSRP, QoS, and High Availability.
- Hands-on experience with enterprise networking platforms such as Cisco, Juniper, Arista, NVIDIA, DELL and Data Center switching and routing technologies.
- Strong experience supporting Firewalls, VPNs, Load Balancers (F5 or equivalent), DNS, DHCP, NTP, SNMP, Syslog, and network security operations.
- Extensive experience in COLO and Cloud Data Center environments, including Azure, AWS, and hybrid cloud networking.
- Strong knowledge of Linux/Windows administration, virtualization platforms, storage, backup, and enterprise infrastructure services.
- Experience with enterprise Network Monitoring, Observability, and AIOps platforms, including SolarWinds, Prometheus, Grafana, Splunk/ELK, Checkmk, NetQ/UFM, or equivalent tools.
- Experience with ITSM platforms such as Zoho, ServiceNow, BMC Remedy, or Jira Service Management, with a strong understanding of Incident, Problem, Change, and Major Incident Management aligned to ITIL.
- Experience with automation and Infrastructure as Code (IaC) using Python, PowerShell, Ansible, Terraform, or equivalent technologies.
- Knowledge of Kubernetes, container platforms, and cloud-native infrastructure operations is highly desirable.
- Strong understanding of Industry security standards like ISO 27001, PCI-DSS, SOC-1&2
- Strong analytical, Root Cause Analysis (RCA), performance optimization, and capacity planning skills.
- Proven leadership, stakeholder management, technical mentoring, and cross-functional collaboration skills.
- ITIL Foundation/Intermediate certification.
- CCNP Enterprise or CCIE (preferred).
- Cloud certifications such as Microsoft Azure Administrator/Architect, AWS SysOps Administrator/Solutions Architect, or equivalent.
- Vendor certifications from Juniper (JNCIP/JNCIE), Arista (ACE), Red Hat (RHCSA/RHCE), VMware (VCP), or equivalent is an added advantage.
- Experience supporting AI GPU clusters, NVIDIA InfiniBand/Ethernet fabrics, high-performance storage, and large-scale Kubernetes environments.
- Familiarity with GPU monitoring, fabric telemetry, distributed storage, and observability platforms used in AI/HPC environments.
- Experience in automation and operational optimization for hyperscale or AI Cloud Data Centers.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search