Back to search
iSoftStone Linkedin · Posted 20d ago

Site Reliability Engineer

WP. Kuala Lumpur

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

About Us:

A leading global technology group, renowned for its extensive ecosystem of digital services and platforms. With a strong presence in cloud computing, mobile gaming, social media, and enterprise solutions, the organization supports millions of users and businesses worldwide. It emphasizes innovation, scalability, and security, making it a key player in driving digital transformation across various industries.


Job Responsibilities:

  • Participate in the architecture design, deployment, and maintenance of overseas game application platforms.
  • Ensure the high availability, reliability, and scalability of game platform services, particularly account and data storage services.
  • Monitor and maintain production environments, proactively identifying and resolving performance, availability, and reliability issues.
  • Design, optimize, and maintain monitoring and observability solutions to improve platform visibility and operational efficiency.
  • Support incident troubleshooting, root cause analysis (RCA), and preventive measures to minimize service disruptions.
  • Automate routine operational tasks and continuously improve deployment, monitoring, and maintenance processes.
  • Maintain technical documentation, operational procedures, and troubleshooting guidelines.


Job Requirements:

  • Bachelor's degree or above in Computer Science, Software Engineering, Information Technology, or a related technical field.
  • Minimum 1 year of hands-on experience in SRE, DevOps, Cloud Engineering, or Platform Operations. Experience in the gaming industry is a strong advantage.
  • Strong knowledge of Unix/Linux operating systems, with hands-on experience in system troubleshooting.
  • Practical experience with Shell and/or Python scripting for automation and operational tasks.
  • Hands-on experience managing public cloud platforms, such as AWS or GCP.
  • Solid experience with Kubernetes (K8s) and its ecosystem, including containerized application deployment and operations.
  • Working experience with MySQL, Redis, or related database technologies.
  • Understanding of monitoring and observability, with experience using tools such as Prometheus, Grafana, ELK/EFK, or similar technologies is an advantage.
  • Good English and Chinese (Mandarin) communication skills are required, as the role involves collaboration with global team.
  • Software development experience in Python, Go, or other programming languages is a plus point.
Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search