Back to search
Tap Growth ai Linkedin · Posted 9d ago

Site Reliability Engineer

Shanghai

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

🌟 We're Hiring: Site Reliability Engineer! 🌟

We are seeking a Site Reliability Engineer (SRE) to ensure the stability, availability, and performance of mission-critical applications and cloud infrastructure. The ideal candidate has strong expertise in incident management, system reliability, cloud operations, and automation, with hands-on experience supporting production environments on Alibaba Cloud.

📍 Location: Shanghai, China
⏰ Work Mode: Work from Office
💼 Role: Site Reliability Engineer

What You'll Do:

  • Monitor, analyze, troubleshoot, and resolve system instability, application crashes, and production incidents to ensure high system availability and reliability.
  • Perform root cause analysis (RCA), implement permanent fixes, and drive continuous improvements to reduce recurring issues and improve platform resilience.
  • Build and maintain monitoring, alerting, logging, and observability solutions to proactively identify and address system health issues.
  • Collaborate with software engineering, DevOps, infrastructure, and platform teams to improve application performance, scalability, and operational efficiency.
  • Automate operational processes, deployments, and infrastructure management while promoting SRE best practices, reliability engineering, and operational excellence.

General Background

  • Bachelor's Degree in Computer Science, Information Technology, Software Engineering, Engineering, or a related discipline.
  • Minimum 7 years of experience in Site Reliability Engineering (SRE), DevOps, Cloud Operations, Infrastructure Engineering, or Production Support.
  • Experience supporting high-availability, cloud-native, or large-scale enterprise production environments.
  • Experience working within Agile or DevOps environments with cross-functional engineering teams.

Mandatory Skills

  • Hands-on experience with Alibaba Cloud services, including cloud infrastructure, networking, compute, storage, monitoring, and security.
  • Strong knowledge of Linux systems administration, incident management, root cause analysis, performance tuning, system monitoring, and troubleshooting production issues.
  • Experience with automation, scripting (Shell, Python, or Bash), CI/CD pipelines, container technologies (Docker, Kubernetes), and observability tools for monitoring and logging.

Nice-to-Have Skills

  • Experience with Infrastructure-as-Code (Terraform, Ansible, or similar automation tools).
  • Familiarity with distributed systems, microservices, event-driven architectures, and cloud-native application deployments.
  • Alibaba Cloud certifications or other cloud platform certifications, with experience supporting 24/7 production operations and Service Level Objectives (SLOs).

Ready to make an impact? 🚀 Apply now and let's grow together!

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search