CloudOps Engineer (L3)
Indexed description
Roles & Responsibilities
Infrastructure Architecture & Engineering
- Design, implement, administer, and optimize enterprise infrastructure across Windows, Linux, virtualization, Kubernetes, storage, backup, cloud, and GPU platforms.
- Define technical standards, best practices, and operational procedures to improve infrastructure reliability and scalability.
- Lead the resolution of complex and business-critical infrastructure incidents.
- Perform detailed Root Cause Analysis (RCA), incident chronology and implement permanent corrective and preventive actions.
- Provide technical leadership during Major Incident Management and disaster recovery scenarios.
- Act as the highest escalation point for L1 and L2 support teams.
- Design and administer enterprise Active Directory, Group Policy, DNS, DHCP, Certificate Services, Federation Services, and Windows Server environments.
- Lead server lifecycle management, security hardening, patch strategy, and performance optimization.
- Administer enterprise Linux platforms, including kernel tuning, automation, clustering, security hardening, high availability, and performance optimization.
- Troubleshoot complex OS, application, and platform-related issues.
- Design, administer, and optimize VMware, Hyper-V, and Kubernetes/OpenShift environments.
- Manage cluster architecture, capacity planning, workload optimization, upgrades, and platform lifecycle management.
- Design and administer enterprise storage and backup solutions.
- Lead disaster recovery planning, backup architecture, replication, capacity planning, and business continuity initiatives.
- Support Azure, AWS, or Google Cloud infrastructure, hybrid networking, and cloud platform administration.
- Design, administer, and optimize GPU infrastructure supporting AI, Machine Learning, and High-Performance Computing (HPC) workloads.
- Manage GPU resource allocation, driver lifecycle, firmware updates, monitoring, and performance optimization.
- Develop infrastructure automation using PowerShell, Bash, Python, Ansible, Terraform, or similar technologies.
- Drive Infrastructure-as-Code (IaC), configuration management, and operational automation initiatives.
- Review and implement complex infrastructure changes.
- Lead infrastructure upgrades, migrations, platform refresh, and technology transformation projects.
- Perform capacity planning and infrastructure performance optimization.
- Mentor L1 and L2 engineers through technical guidance and knowledge sharing.
- Develop technical standards, runbooks, SOPs, and operational documentation.
- Collaborate with architects, engineering teams, OEMs, and vendors on strategic initiatives.
Relevant Experience
- 6 to 10+ years in Enterprise Infrastructure Engineering, Managed Services, Data Center Operations, Cloud Infrastructure, or Platform Engineering.
- Proven experience supporting business-critical production environments with 24×7 operations.
- Hands-on experience with enterprise monitoring tools (SolarWinds, SCOM, Zabbix, Nagios, PRTG, Prometheus, Grafana) and ITSM platforms such as ServiceNow, ManageEngine, BMC Remedy, or Jira Service Management. Expertise in virtualization (VMware, Hyper-V, Kubernetes, OpenShift), enterprise storage, backup solutions, and hybrid cloud platforms (Azure, AWS, GCP). Strong exposure to GPU infrastructure for AI/ML and HPC workloads, Infrastructure-as-Code (Terraform, Ansible), container platforms, CI/CD pipelines, and enterprise security technologies.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search