Back to search
Yum! Brands Linkedin · Posted 10d ago

Site Reliability Engineering Manager

Ho Chi Minh City

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Key Responsibilities


Team Leadership and Development

  • Lead and develop a team of Site Reliability Engineers (Level 6-7), initially 3 direct reports and growing as the Vietnam site consolidates, owning performance, coaching, retention, and day-to-day execution
  • Build individual development plans that grow Level 6 engineers toward independent Level 7 scope
  • Establish a high-ownership culture where engineers are accountable for outcomes, not tasks
  • Run team rituals: 1:1s, shift retrospectives, and development check-ins


Follow-the-Sun Operations

  • Own the Vietnam shift within GRE’s global follow-the-sun coverage model, including schedule design, coverage planning, and holiday/leave management
  • Ensure clean, structured handoffs to and from US and India teams, with clear ownership transfer on open incidents and in-flight work
  • Maintain shift readiness: runbooks current, alerts actionable, escalation paths clear
  • Serve as escalation point for the Vietnam shift during complex or high-severity incidents


Reliability Practice Execution

  • Own production reliability outcomes for the markets, platforms, and services within the Vietnam team’s scope
  • Drive SRE operating standards within the team: incident response rigor, SLO ownership, runbook quality, and post-incident follow-through
  • Ensure monitoring coverage, dashboards, and alerting remain accurate and effective across owned services
  • Enforce GRE-wide reliability standards, ensuring the Vietnam practice operates in alignment with the global model


Platform Engineering, Automation, and AI

  • Own the Vietnam team’s contribution to GRE’s reliability modernization roadmap: auto-healing, auto-remediation, and self-service issue mitigation
  • Treat automation as a delivery commitment, not a byproduct: plan, prioritize, and track toil-reduction and self-service work alongside operational coverage
  • Ensure the team builds with platform engineering practices: infrastructure as code, runbooks as code, GitOps workflows, and reusable tooling over one-off fixes
  • Champion AI-first engineering practices within the team, ensuring engineers develop with and through modern AI tooling
  • Partner with GRE’s Foundations and Intelligence pillars on observability, signal detection, and automated response infrastructure


Stakeholder Engagement

  • Represent the Vietnam team in East GRE planning and reliability forums
  • Communicate reliability outcomes, coverage status, and team health clearly to GRE leadership
  • Partner with the Associate Director on headcount planning, team scope, and Vietnam-specific delivery


Mandatory Skills


  • 5+ years in site reliability engineering, DevOps, infrastructure, or production operations roles
  • 1+ years of people management experience, or 2+ years as a senior technical lead with demonstrated coaching and delivery ownership
  • Hands-on credibility across incident response, observability, and automation, with the technical depth to guide Level 6-7 engineers
  • Experience operating in shift-based, on-call, or follow-the-sun coverage models
  • Working knowledge of at least one major cloud provider (AWS preferred) and modern observability tooling (e.g., Datadog, Prometheus, Grafana)
  • Proficiency in at least one scripting or programming language sufficient to review and guide automation work
  • Understanding of SLI/SLO frameworks and reliability engineering fundamentals
  • Strong written and verbal English communication skills for cross-region collaboration with US and India teams
  • Experience building or standing up a new team, site, or shift operation
  • Experience managing engineers across the early-to-mid career range with a track record of promotions or level progression
  • Kubernetes, container orchestration, and infrastructure as code experience (e.g., Terraform)
  • Exposure to platform engineering and internal developer platform concepts: self-service tooling, developer portals (e.g., Port, Backstage), GitOps
  • Experience in multi-region or globally distributed team models
  • Relevant certifications (AWS, CKA, or similar)
Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search