Back to search
HFG Insurance Recruitment Linkedin · Posted 5d ago

Senior / Lead Platform Reliability Engineer (PRE)

Cyberjaya

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Hiring: Senior / Lead Platform Reliability Engineer (PRE)

Location: Cyberjaya, Malaysia | Infrastructure / Platform Engineering

We are looking for experienced Platform Reliability Engineers to join an enterprise infrastructure team responsible for engineering, operating and continuously improving a highly available internal container platform.

This is an excellent opportunity for engineers with strong Tanzu / Kubernetes, automation, CI/CD, networking and platform reliability experience who enjoy working on large-scale enterprise environments.

Key Responsibilities

  • Engineer, operate and maintain enterprise-grade container platforms and supporting infrastructure, with a strong focus on reliability, resiliency, security and performance.
  • Work extensively with Broadcom VMware Tanzu and Kubernetes-based container orchestration platforms.
  • Perform platform resource provisioning, capacity planning, monitoring, performance optimization and reliability improvements.
  • Provide L2/L3 production support, including troubleshooting complex platform, infrastructure, application and networking issues.
  • Participate in a 24/7 on-call rotation and respond to critical production alerts and incidents.
  • Lead or support major incident management, including troubleshooting, vendor coordination, immediate remediation, root cause analysis and long-term corrective actions.
  • Design and implement automation using Ansible, Python and Bash to reduce manual processes and operational errors.
  • Develop and maintain Helm charts and Helm repositories.
  • Work with Docker, Kubernetes, CI/CD and source-control technologies, including Bitbucket and related DevOps tooling.
  • Manage platform upgrades, patching, product updates and buildpack management.
  • Monitor and optimize platform reliability using Grafana and Dynatrace, with a focus on SLOs, SLIs and SLAs.
  • Troubleshoot and resolve complex infrastructure, networking and container platform issues.
  • Review security advisories and ensure timely remediation and updates across the container platform.
  • Work closely with application, infrastructure, security, network and other technical teams to deliver reliable enterprise solutions.
  • For the Lead level, provide technical leadership to an existing PRE team, manage on-call resources, coordinate platform upgrades and deployment activities, and oversee ServiceNow/Jira queues and SLA delivery.

Technical Requirements

  • Bachelor's or Master's degree in Computer Science, IT or a related discipline.
  • Strong hands-on experience with Tanzu Application Service (TAS), Tanzu Kubernetes Grid Integrated Edition (TKGI), or Kubernetes-based platforms.
  • Strong experience with Ansible, Python and Bash scripting.
  • Hands-on experience developing and maintaining Helm charts and Helm repositories.
  • Experience with NSX-T and integration with Tanzu/Kubernetes environments.
  • Strong understanding of Kubernetes, Docker and container technologies.
  • Experience with CI/CD pipelines, SCM and DevOps practices.
  • Experience with platform upgrades, patching and buildpack management.
  • Strong troubleshooting capabilities across infrastructure, networking and container platforms.
  • Experience with Grafana and/or Dynatrace, including monitoring SLOs, SLIs and SLAs.
  • Familiarity with Bamboo, Bitbucket, Nexus, Jira and Confluence.
  • Strong understanding of reliability engineering, scalability, performance optimization and enterprise platform architecture.
  • Excellent stakeholder management, communication and documentation skills.

Senior Platform Reliability Engineer

  • 5–7 years of overall IT experience.
  • 3–5 years of hands-on experience in Platform Reliability Engineering or Site Reliability Engineering.
  • 3+ years of automation experience using Ansible, Python and Bash.
  • 3+ years of Helm experience.
  • 3+ years of NSX-T experience.
  • Experience operating in high-demand, fast-paced production environments.

Lead Platform Reliability Engineer

  • 7–10+ years of hands-on experience with container orchestration platforms.
  • 5+ years of automation experience using Ansible, Python and Bash.
  • 5+ years of Helm experience.
  • 3+ years of NSX-T and Tanzu integration experience.
  • Strong experience in enterprise platform architecture and reliability engineering.
  • Proven experience leading technical teams and managing production operations.
  • Experience owning L3 support, major incidents, platform upgrades and operational delivery.

Certifications

One or more of the following is highly desirable:

  • Certified Kubernetes Administrator (CKA)
  • Certified Kubernetes Application Developer (CKAD)
  • Certified Kubernetes Security Specialist (CKS)

Ideal Candidate

We are looking for someone who is technically hands-on, proactive and comfortable taking ownership of production platforms. You should enjoy solving complex infrastructure and networking problems, automating repetitive processes and continuously improving platform reliability.

Candidates with strong Tanzu/Kubernetes, NSX-T, Ansible, Helm and CI/CD experience are encouraged to apply.

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search