System Engineer (Nutanix and Red Hat)
Indexed description
As a System Engineer dedicated to Day 2 Operations, you will focus on maintaining, optimizing, and ensuring the absolute stability of our enterprise hybrid infrastructure. This role is well suited to someone who excels at the "post-deployment" lifecycle—keeping business-critical Nutanix hyperconverged infrastructure (HCI) systems and Red Hat Enterprise Linux (RHEL) platforms secure, performant, resilient, and continuously automated.
Responsibilities
- Day 2 Operations & Lifecycle Management: Own the monitoring, proactive health-checking, performance tuning, and capacity management of enterprise Nutanix HCI clusters (AOS/AHV) and RHEL environments
- Patching & Compliance: Manage routine and emergency patching, kernel upgrades, and firmware updates across Nutanix blocks and RHEL estates while adhering strictly to government security baselines and minimizing system downtime
- Infrastructure Automation: Write and maintain Ansible playbooks and Bash scripts to automate repetitive system administration tasks, configurations, and configuration drift management
- Incident & Problem Management: Serve as the technical escalation point for complex infrastructure incidents, conducting deep-dive root cause analysis (RCA) and implementing permanent fixes
- Backup & Disaster Recovery: Regularly manage, validate, and test disaster recovery (DR) procedures and backup replication policies across our production nodes
- Collaboration: Work closely with application developers, security operators, and infrastructure architects to ensure platform changes cleanly align with Day 2 operational stability
- 3-5 years of hands-on experience in enterprise systems engineering, with a heavy emphasis on production infrastructure operations (Day 2 support)
- Strong competency managing Nutanix environments, including Nutanix AHV/ESXi, Prism Central, and lifecycle management (LCM)
- Deep expertise in Red Hat Enterprise Linux (RHEL) administration, covering storage management (LVM), performance tuning, security hardening, and OS troubleshooting
- Hands-on experience with Ansible for system automation, configuration management, and patching workflows
- Solid understanding of core enterprise infrastructure concepts: networking (VLANs, routing), storage protocols, and security principles
- Familiarity with enterprise monitoring pipelines (e.g., Nutanix Pulse, Prometheus/Grafana, or Syslog)
- Nutanix Certified Professional (NCP) or Red Hat Certified System Administrator / Engineer (RHCSA / RHCE)
- Experience operating infrastructure within a highly regulated environment or Singapore public sector context
- Basic exposure to container platforms like Red Hat OpenShift or vanilla Kubernetes
- Willingness to participate in a scheduled on-call rotation for high-severity production incident resolution
- We promote a learning culture and encourage you to grow and learn with training budgets and certification sponsorships
- Annual Leave Benefits with additional perks such as Family Care and Birthday Leave
- Contract Staff enjoys the same comprehensive benefits as Permanent Employees
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search