Infrastructure Engineer
Indexed description
Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud
Hyper- V SRE ( Senior Infrastructure Engineer — Hyper-V
)
Mandatory Skills for Hyper-V SRE Production Suppo
rtMicrosoft Hyper-V Administration (Deployment, Troubleshooting, Optimizatio
n)Hyper-V Failover Clustering & High Availabili
tyWindows Server 2016/2019/2022 Administration & OS patchi
ngStorage Spaces Direct (S2D), CSV, SAN/NAS & Stora
geSite Reliability Engineering (SRE) Principles, SLI/SLO, Reliabili
tyPowerShell Scripting & Automati
onDisaster Recovery, Backup, Hyper-V Replica & Business Continui
tySCVMM (System Center Virtual Machine Manage
r
) Desired Skills for Hyper-V SRE Production Suppo
rtAzure Monitor / SCOM / Prometheus / Grafana / Splu
nkVeeam / Altaro Backup Solutio
nsVDI Infrastructu
reInfrastructure as Code (Ia
C)Hybrid Cloud / Private Clo
udITIL (Incident, Change, Problem Managemen
t)
Over
viewWe are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern Site Reliability Engineering practices to improve platform reliability, scalability, performance, automation, and operational excelle
nce.The ideal candidate will be responsible for ensuring platform reliability through infrastructure automation, OS upgrades, proactive monitoring, incident response, capacity planning, and continuous service improvement while collaborating with infrastructure, security, and application te
ams.Key Responsibili
- tiesOperate enterprise-scale private cloud infrastructure built on Microsoft Hype
- r-V.Optimize, and support highly available VDI environments on Hype
- r-V.Improve platform reliability, availability, scalability, and resiliency by applying SRE principles and engineering best practi
- ces.Disaster recovery, backup, patch management, and business continuity strateg
- ies.Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational metrics for critical infrastructure servi
- ces.Automate infrastructure provisioning, configuration management, and operational workflows using PowerShell and Infrastructure as Code (IaC) principles wherever applica
- ble.Manage Hyper-V Failover Clusters, host lifecycle, storage, networking, and capacity to ensure high availability and business continu
- ity.Develop proactive monitoring, alerting, logging, and observability capabilities to detect and prevent service degradat
- ion.Lead incident response for infrastructure-related outages, perform root cause analysis (RCA), and implement preventive actions through post-incident revi
- ews.Perform capacity planning, performance tuning, and resource optimization across Hyper-V clusters and VDI platfo
- rms.Support infrastructure migration initiatives including P2V, V2V, workload modernization, and private cloud transformati
- ons.Collaborate closely with Security, Networking, Platform Engineering, and Application teams to improve platform reliability and operational efficie
- ncy.Develop and maintain technical documentation, architecture diagrams, operational runbooks, automation scripts, and standard operating procedu
- res.Mentor junior engineers and promote SRE culture, automation, and operational best practices across the t
eam.Required Technical Sk
illsSite Reliability Enginee
- ringStrong understanding of Site Reliability Engineering principles and operational excelle
- nce.Experience with infrastructure reliability, service availability, resiliency, and performance optimizat
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search