Manager - SRE
Indexed description
As an SRE Manager, you will lead a team of site reliability engineers, define best practices, and partner with engineering, product, and infrastructure teams to deliver resilient and scalable systems. You will be responsible for building a culture of reliability, ensuring operational readiness, and mentoring engineers to excel in reliability engineering.
Requirements
- Willingness to work in a 24×7 operational environment
- 10–14 years of overall experience, with at least 3–5 years in a leadership/people management role within SRE, DevOps, or Infrastructure Engineering.
- Strong technical background in system design, distributed systems, and cloud-native architecture.
- Hands-on expertise in automation, containerization (Docker, Kubernetes), CI/CD tools, and configuration management (Ansible, Chef, etc.).
- Proficiency in programming/scripting languages (Python, Bash, Go, etc.) for automation and tooling.
- Deep understanding of observability platforms (Grafana, Splunk, Dynatrace, Prometheus, etc.) and implementing monitoring strategies.
- Proven experience with incident and problem management, including driving postmortems and continuous improvement.
- Familiar with SLI, SLO, SLA, and Error Budget frameworks and applying them across multiple services.
- Strong knowledge of networking, Unix/Linux systems, cloud platforms (AWS, GCP, or Azure), and databases.
- Ability to drive reliability goals across multiple teams while balancing delivery speed and stability.
- Excellent communication, leadership, and stakeholder management skills.
- Lead and mentor a team of SREs, fostering a culture of accountability, collaboration, and technical excellence.
- Design, own, and continuously improve the team's 24×7 on-call rotation and escalation matrix, ensuring sustainable coverage while personally stepping in as the point of escalation for critical production issues or as & when needed as a owner
- Drive the design, implementation, and adoption of scalable, reliable, and secure infrastructure solutions.
- Partner with engineering and product teams to embed reliability best practices into the development lifecycle.
- Define and track reliability metrics (SLIs/SLOs/SLAs) and ensure adherence across services.
- Oversee incident response, root cause analysis, and implement learnings to minimize recurrence.
- Build automation frameworks to reduce operational toil and improve developer productivity.
- Own capacity planning, disaster recovery strategies, and performance optimization of systems.
- Act as the point of escalation for critical production issues and provide leadership during incidents.
- Establish long-term strategies for reliability, scalability, and cost efficiency of the platform.
- Represent the SRE function with senior leadership, providing updates on reliability posture and improvement plans.
In retail stores, our gStore end-to-end store execution and retail management solution supports omnichannel fulfillment, real-time replenishment, intelligent workforce tasking and more. Using real-time overhead RFID technology, the platform increases inventory accuracy up to 99%, doubles staff productivity, and enables an engaging, seamless in-store experience.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search