Back to search
Yum! Brands Linkedin · Posted 9d ago

Site Reliability Engineer

Ho Chi Minh City

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Key Responsibilities


Incident Management and Reliability

  • Lead or coordinate incident response for assigned markets and services, with independent ownership of routine incidents and supported ownership of complex multi-system incidents
  • Facilitate communication during incidents and drive timely resolution, including clean shift handoffs across global time zones
  • Contribute to post-incident reviews and root cause analysis activities
  • Ensure corrective and preventive actions are identified, tracked, and completed for owned services


Monitoring, Alerting, and Observability

  • Implement and continuously optimize monitoring, logging, alerting, and tracing solutions
  • Develop meaningful alerts based on service behavior, customer impact, and business priorities
  • Build and maintain dashboards that provide actionable insights into system performance and reliability


Platform and Market Owners

  • Own day-to-day SRE responsibilities for one or more assigned markets, platforms, or services
  • Ensure monitoring coverage, dashboards, runbooks, and alerting configurations remain accurate, effective, and up to date
  • Assess platform health, identify reliability risks, and raise improvements before incidents occur
  • Partner with engineering teams to ensure new features and services meet reliability requirements before production release
  • Track reliability metrics for owned services, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets


Platform Engineering, Automation and AI

  • Develop and maintain tools, scripts, and automation that reduce manual effort and improve operational efficiency
  • Identify and eliminate repetitive tasks through automation, with a bias toward self-service capabilities over ticket queues.
  • Contribute to auto-healing and auto-remediation capabilities: detection rules, remediation runbooks as code, and automated response workflows.
  • Build and extend internal platform tooling using infrastructure as code and CI/CD pipelines rather than manual configuration
  • Apply team best practices for the responsible use of AI within SRE workflows, including AI-assisted diagnosis, triage, and remediation


Mandatory Skills


  • 2+ years of experience in SRE, DevOps, production support, or infrastructure engineering roles.
  • Hands-on experience with monitoring and observability tooling (e.g., Datadog, Prometheus, Grafana, CloudWatch, or similar)
  • Working knowledge of at least one major cloud provider (AWS preferred)
  • Proficiency in at least one scripting or programming language (e.g., Python, Bash, Go) for automation, with demonstrated examples of automating away manual operational work.
  • Experience participating in incident response and on-call or shift-based operations.
  • Understanding of SLI/SLO concepts and reliability engineering fundamentals.
  • Ability to work follow-the-sun shift rotations, including structured handoffs with teams in other regions.
  • Strong written and verbal English communication skills for cross-region collaboration.
  • Experience with Kubernetes, Docker, and container orchestration in production.
  • Infrastructure as code experience (e.g., Terraform, CloudFormation).
  • Experience building or contributing to auto-remediation or event-driven automation workflows.
  • Experience supporting distributed systems across multiple markets or regions.
  • Relevant certifications (AWS, CKA, or similar)
Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search