Site Reliability Engineer (SRE)
Indexed description
🌟 We're Hiring: Site Reliability Engineer (SRE)! 🌟
We are looking for aSite Reliability Engineer (SRE)with hands-on experience inAlibaba Cloud (AliCloud)to support and maintain reliable, scalable, and secure cloud infrastructure. The ideal candidate will work closely with development and operations teams to improve system performance, automate processes, and ensure high service availability.
Key Responsibilities- Manage and support infrastructure hosted on Alibaba Cloud (AliCloud).
- Monitor system performance, availability, and reliability.
- Implement automation using scripting and Infrastructure as Code (IaC) tools.
- Support Kubernetes and containerized environments.
- Troubleshoot incidents, perform root cause analysis, and implement preventive measures.
- Maintain monitoring, alerting, and observability solutions.
- Collaborate with development and security teams to improve platform stability and performance.
- Participate in on-call support and incident response activities.
- 8+ years of experience in SRE, DevOps, Cloud Engineering, or related roles.
- Hands-on experience withAlibaba Cloud (AliCloud)services.
- Strong Linux administration skills.
- Experience with Kubernetes, Docker, and cloud-native technologies.
- Knowledge of Terraform, Ansible, or other automation tools.
- Experience with monitoring tools such as Prometheus, Grafana, ELK, or CloudMonitor.
- Scripting skills in Python, Bash, or Shell.
- Good understanding of networking, security, and high-availability architectures.
- Strong troubleshooting and problem-solving skills.
- Alibaba Cloud certifications are a plus.
- Experience with CI/CD tools such as Jenkins, GitLab CI/CD, or GitHub Actions.
- Exposure to AWS, Azure, or GCP is an advantage.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search