Back to search
Pipsfortunes Linkedin · Posted yesterday

Azure SRE Lead – Azure DevOps / Kubernete - Only W2

Atlanta, Georgia, United States

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Job Description

We are seeking an experienced Azure SRE Lead with strong expertise in Site Reliability Engineering, Azure DevOps, Kubernetes, observability, and cloud infrastructure. The ideal candidate will have 5+ years of hands-on SRE experience and a proven track record supporting large-scale distributed applications and production environments.

The candidate will be responsible for improving system reliability, monitoring and observability, automation, incident management, performance, and production operations across cloud-based environments.

Key Responsibilities
  • Lead Site Reliability Engineering initiatives for large-scale distributed applications.
  • Design and implement reliable, scalable, and highly available cloud infrastructure.
  • Build and maintain Azure DevOps CI/CD pipelines and deployment automation.
  • Manage and support Kubernetes and Docker containerized environments.
  • Implement observability solutions covering logs, metrics, traces, and application telemetry.
  • Develop dashboards, alerts, and monitoring solutions using tools such as Splunk, Datadog, Dynatrace, Grafana, Prometheus, and OpenTelemetry.
  • Collect, analyze, and troubleshoot telemetry data to identify system performance and reliability issues.
  • Lead production incident response, root cause analysis, and remediation activities.
  • Perform performance tuning and capacity planning for distributed applications.
  • Implement Infrastructure as Code and automation for cloud environments.
  • Develop automation scripts using Python, Bash, or Go.
  • Establish SRE best practices around reliability, availability, scalability, and operational efficiency.
  • Collaborate with development, DevOps, cloud, and infrastructure teams to improve application reliability.
  • Support continuous improvement of production environments and operational processes.
Must-Have Skills
  • 5+ years of hands-on experience in Site Reliability Engineering (SRE).
  • Strong experience with Azure DevOps.
  • Strong hands-on experience with Kubernetes.
  • Experience with Docker and containerized applications.
  • Strong knowledge of observability and monitoring.
  • Experience with Splunk, Datadog, Dynatrace, Grafana, Prometheus, or OpenTelemetry.
  • Experience with logging, metrics, tracing, telemetry collection, and dashboard development.
  • Strong experience with cloud platforms, preferably Microsoft Azure.
  • Hands-on experience with CI/CD pipelines and automation.
  • Experience with Infrastructure as Code.
  • Strong troubleshooting, performance tuning, and incident management skills.
  • Proficiency in Python, Bash, or Go.
  • Experience supporting large-scale distributed applications in production.
Nice-to-Have Skills
  • Microsoft Fabric
  • Experience with advanced Azure monitoring and observability services.
  • Experience with multi-cloud environments such as AWS or GCP.
  • Experience implementing SRE frameworks and reliability metrics/SLIs/SLOs.
  • Experience with automated remediation and self-healing infrastructure.
Qualifications
  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field preferred.
  • 8–10 years of overall IT experience with strong SRE/DevOps experience.
  • Strong communication, leadership, troubleshooting, and problem-solving skills.
  • Ability to work in a hybrid environment in Atlanta, GA (2 days per week onsite)


Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search