Azure SRE Lead – Azure DevOps / Kubernete - Only W2
Indexed description
We are seeking an experienced Azure SRE Lead with strong expertise in Site Reliability Engineering, Azure DevOps, Kubernetes, observability, and cloud infrastructure. The ideal candidate will have 5+ years of hands-on SRE experience and a proven track record supporting large-scale distributed applications and production environments.
The candidate will be responsible for improving system reliability, monitoring and observability, automation, incident management, performance, and production operations across cloud-based environments.
Key Responsibilities- Lead Site Reliability Engineering initiatives for large-scale distributed applications.
- Design and implement reliable, scalable, and highly available cloud infrastructure.
- Build and maintain Azure DevOps CI/CD pipelines and deployment automation.
- Manage and support Kubernetes and Docker containerized environments.
- Implement observability solutions covering logs, metrics, traces, and application telemetry.
- Develop dashboards, alerts, and monitoring solutions using tools such as Splunk, Datadog, Dynatrace, Grafana, Prometheus, and OpenTelemetry.
- Collect, analyze, and troubleshoot telemetry data to identify system performance and reliability issues.
- Lead production incident response, root cause analysis, and remediation activities.
- Perform performance tuning and capacity planning for distributed applications.
- Implement Infrastructure as Code and automation for cloud environments.
- Develop automation scripts using Python, Bash, or Go.
- Establish SRE best practices around reliability, availability, scalability, and operational efficiency.
- Collaborate with development, DevOps, cloud, and infrastructure teams to improve application reliability.
- Support continuous improvement of production environments and operational processes.
- 5+ years of hands-on experience in Site Reliability Engineering (SRE).
- Strong experience with Azure DevOps.
- Strong hands-on experience with Kubernetes.
- Experience with Docker and containerized applications.
- Strong knowledge of observability and monitoring.
- Experience with Splunk, Datadog, Dynatrace, Grafana, Prometheus, or OpenTelemetry.
- Experience with logging, metrics, tracing, telemetry collection, and dashboard development.
- Strong experience with cloud platforms, preferably Microsoft Azure.
- Hands-on experience with CI/CD pipelines and automation.
- Experience with Infrastructure as Code.
- Strong troubleshooting, performance tuning, and incident management skills.
- Proficiency in Python, Bash, or Go.
- Experience supporting large-scale distributed applications in production.
- Microsoft Fabric
- Experience with advanced Azure monitoring and observability services.
- Experience with multi-cloud environments such as AWS or GCP.
- Experience implementing SRE frameworks and reliability metrics/SLIs/SLOs.
- Experience with automated remediation and self-healing infrastructure.
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field preferred.
- 8–10 years of overall IT experience with strong SRE/DevOps experience.
- Strong communication, leadership, troubleshooting, and problem-solving skills.
- Ability to work in a hybrid environment in Atlanta, GA (2 days per week onsite)
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search