Senior / Lead Platform Reliability Engineer (PRE)
Indexed description
Hiring: Senior / Lead Platform Reliability Engineer (PRE)
Location: Cyberjaya, Malaysia | Infrastructure / Platform Engineering
We are looking for experienced Platform Reliability Engineers to join an enterprise infrastructure team responsible for engineering, operating and continuously improving a highly available internal container platform.
This is an excellent opportunity for engineers with strong Tanzu / Kubernetes, automation, CI/CD, networking and platform reliability experience who enjoy working on large-scale enterprise environments.
Key Responsibilities
- Engineer, operate and maintain enterprise-grade container platforms and supporting infrastructure, with a strong focus on reliability, resiliency, security and performance.
- Work extensively with Broadcom VMware Tanzu and Kubernetes-based container orchestration platforms.
- Perform platform resource provisioning, capacity planning, monitoring, performance optimization and reliability improvements.
- Provide L2/L3 production support, including troubleshooting complex platform, infrastructure, application and networking issues.
- Participate in a 24/7 on-call rotation and respond to critical production alerts and incidents.
- Lead or support major incident management, including troubleshooting, vendor coordination, immediate remediation, root cause analysis and long-term corrective actions.
- Design and implement automation using Ansible, Python and Bash to reduce manual processes and operational errors.
- Develop and maintain Helm charts and Helm repositories.
- Work with Docker, Kubernetes, CI/CD and source-control technologies, including Bitbucket and related DevOps tooling.
- Manage platform upgrades, patching, product updates and buildpack management.
- Monitor and optimize platform reliability using Grafana and Dynatrace, with a focus on SLOs, SLIs and SLAs.
- Troubleshoot and resolve complex infrastructure, networking and container platform issues.
- Review security advisories and ensure timely remediation and updates across the container platform.
- Work closely with application, infrastructure, security, network and other technical teams to deliver reliable enterprise solutions.
- For the Lead level, provide technical leadership to an existing PRE team, manage on-call resources, coordinate platform upgrades and deployment activities, and oversee ServiceNow/Jira queues and SLA delivery.
Technical Requirements
- Bachelor's or Master's degree in Computer Science, IT or a related discipline.
- Strong hands-on experience with Tanzu Application Service (TAS), Tanzu Kubernetes Grid Integrated Edition (TKGI), or Kubernetes-based platforms.
- Strong experience with Ansible, Python and Bash scripting.
- Hands-on experience developing and maintaining Helm charts and Helm repositories.
- Experience with NSX-T and integration with Tanzu/Kubernetes environments.
- Strong understanding of Kubernetes, Docker and container technologies.
- Experience with CI/CD pipelines, SCM and DevOps practices.
- Experience with platform upgrades, patching and buildpack management.
- Strong troubleshooting capabilities across infrastructure, networking and container platforms.
- Experience with Grafana and/or Dynatrace, including monitoring SLOs, SLIs and SLAs.
- Familiarity with Bamboo, Bitbucket, Nexus, Jira and Confluence.
- Strong understanding of reliability engineering, scalability, performance optimization and enterprise platform architecture.
- Excellent stakeholder management, communication and documentation skills.
Senior Platform Reliability Engineer
- 5–7 years of overall IT experience.
- 3–5 years of hands-on experience in Platform Reliability Engineering or Site Reliability Engineering.
- 3+ years of automation experience using Ansible, Python and Bash.
- 3+ years of Helm experience.
- 3+ years of NSX-T experience.
- Experience operating in high-demand, fast-paced production environments.
Lead Platform Reliability Engineer
- 7–10+ years of hands-on experience with container orchestration platforms.
- 5+ years of automation experience using Ansible, Python and Bash.
- 5+ years of Helm experience.
- 3+ years of NSX-T and Tanzu integration experience.
- Strong experience in enterprise platform architecture and reliability engineering.
- Proven experience leading technical teams and managing production operations.
- Experience owning L3 support, major incidents, platform upgrades and operational delivery.
Certifications
One or more of the following is highly desirable:
- Certified Kubernetes Administrator (CKA)
- Certified Kubernetes Application Developer (CKAD)
- Certified Kubernetes Security Specialist (CKS)
Ideal Candidate
We are looking for someone who is technically hands-on, proactive and comfortable taking ownership of production platforms. You should enjoy solving complex infrastructure and networking problems, automating repetitive processes and continuously improving platform reliability.
Candidates with strong Tanzu/Kubernetes, NSX-T, Ansible, Helm and CI/CD experience are encouraged to apply.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search