Site Reliability Engineering (SRE) Manager
Indexed description
- Define and execute SRE strategies that improve system reliability, availability, scalability, and performance.
- Establish and govern Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational health metrics.
- Lead production readiness reviews, disaster recovery testing, resilience assessments, and operational risk mitigation activities.
- Drive continuous improvement of application stability, service availability, and customer experience.
- Lead major incident response and escalation management for critical production issues.
- Oversee root cause analysis (RCA) processes and ensure corrective actions are implemented and tracked to completion.
- Drive reduction of recurring incidents through engineering improvements, automation, and proactive monitoring.
- Provide executive-level communication during significant incidents and service disruptions.
- Establish monitoring, alerting, logging, tracing, and observability standards across supported platforms.
- Lead implementation of dashboards and operational metrics that provide visibility into service health and customer impact.
- Drive automation initiatives that reduce manual operational effort, improve recovery times, and increase engineering efficiency.
- Promote Infrastructure as Code (IaC), CI/CD integration, automated remediation, and self-service operational capabilities.
- Partner with Engineering and Infrastructure teams to support cloud-native and hybrid application environments.
- Ensure applications are designed and operated using resilient, scalable, and supportable architectures.
- Support modernization initiatives involving Azure cloud services, containers, APIs, microservices, and platform engineering practices.
- Evaluate vendor platforms and third-party services to ensure reliability and operational readiness.
- Drive adoption of AI and Generative AI capabilities to improve incident response, troubleshooting, observability, and operational efficiency.
- Identify opportunities for intelligent automation, anomaly detection, automated diagnostics, and AI-assisted knowledge management.
- Promote responsible AI adoption aligned with enterprise security, governance, and risk standards.
- Recruit, develop, coach, and retain high-performing Site Reliability Engineers, Production Engineers, Automation Engineers, and Observability Engineers.
- Establish career paths, skill development plans, and succession strategies.
- Foster a culture of ownership, accountability, innovation, collaboration, and continuous learning.
- Manage staffing, performance management, compensation recommendations, and organizational development activities.
- Ensure adherence to enterprise risk, cybersecurity, regulatory, and operational control standards.
- Identify and escalate operational risks impacting critical services or customer experiences.
- Support audits, regulatory reviews, disaster recovery exercises, and operational governance programs.
- Site Reliability Engineering (SRE)
- Production Support
- Observability Engineering
- Incident Management
- Operational Automation
- Cloud Reliability
- Platform Operations
- 10+ years of technology experience with application support, infrastructure, cloud, software engineering, or reliability engineering responsibilities.
- 5+ years of leadership experience managing engineering, operations, or SRE teams.
- Experience managing production systems supporting critical business functions.
- Strong knowledge of Site Reliability Engineering principles, including SLOs, observability, automation, incident management, and operational excellence.
- Experience leading major incident response, root cause analysis, and service restoration efforts.
- Experience with cloud platforms, distributed systems, APIs, and modern application architectures.
- Strong communication, analytical, decision-making, and stakeholder management skills.
- Bachelor's degree in Computer Science, Engineering, Information Technology, or related field.
- Experience leading SRE or Production Engineering organizations.
- Experience with Azure cloud technologies and cloud-native architectures.
- Experience with observability platforms such as Dynatrace, Splunk, Datadog, Grafana, Azure Monitor, or OpenTelemetry.
- Experience with scripting and automation technologies including PowerShell, Python, Bash, and APIs.
- Experience with CI/CD, Infrastructure as Code, DevOps, and Platform Engineering practices.
- Experience implementing operational AI use cases including incident analysis, observability analytics, and automated diagnostics.
- Financial services or other highly regulated industry experience preferred.
- Delivers highly available and resilient customer-facing platforms.
- Uses automation to eliminate operational toil and improve efficiency.
- Reduces mean time to detect (MTTD) and mean time to restore (MTTR).
- Establishes strong observability and operational intelligence capabilities.
- Builds a culture of reliability, accountability, and continuous improvement.
- Successfully integrates AI-assisted operations and automation into support workflows.
- develops high-performing teams that balance reliability, speed, risk management, and customer experience.
Location
Buffalo, New York, United States of America
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search