Site Reliability Engineer
Indexed description
Ready to engineer reliability at massive scale? If you thrive on automating everything, solving high-impact production challenges, and building resilient cloud platforms that power millions of users, this opportunity is for you.
A leading Telecom Giant is currently looking for an experienced Site Reliability Engineer (SRE) to join a join a forward-thinking team, someone who loves tackling complex infrastructure challenges, optimizing cloud-native platforms, and ensuring world-class reliability.
What You'll Be Doing
- Manage and optimize AWS cloud infrastructure and Kubernetes (EKS) environments
- Build and automate infrastructure using Terraform (IaC)
- Design and enhance monitoring, observability, and alerting systems using: Prometheus, Grafana, CloudWatch, ELK Stack, Datadog
- Support and improve CI/CD pipelines and deployment automation
- Troubleshoot production incidents, conduct RCA, and implement reliability improvements
- Partner with engineering teams to deploy and scale cloud-native applications
- Support emerging AI/ML workloads, model-serving platforms, and modern infrastructure
- Develop automation solutions through scripting and Infrastructure as Code
- Participate in production support and on-call rotations
Priority to experts skilled in AWS, Kubernetes/EKS, Terraform, and Prometheus/Grafana to drive the next generation of platform excellence.
Ready to build resilient, scalable cloud platforms for one of the world's leading media and technology companies?
Interested candidates, please send your updated resume or connect with me directly to learn more.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search