AI / High-Performance Computing Engineer
Indexed description
The Digital Division leads the digital transformation of DSO through the master planning and policies, delivering digital capabilities through IT infrastructure, and providing one stop service to corporate and R&D Divisions. The Digital Division will transform the way we work, our workplace, and the capabilities we deliver to the MINDEF/SAF and for the security of Singapore.
In the Integrated Digital Systems, AI HPC Engineers design, build, operate, and optimise the computational workloads and AI inference engines that powers DSO's AI/ML research and helps to enable impactful R&D outcomes by researchers, scientists and engineers.
People are DSO’s greatest asset. You will get to realise your career aspirations and develop your own niche either as a deep technical expert or a leader in the team. With frequent career dialogues and a robust training and development framework, we will provide you with the necessary development tools for you to reach your potential. You will also be recognised and rewarded through competitive remuneration packages and scholarship opportunities.
AI / High-Performance Computing Engineer
In This Role, You Will
In this role, you will:
- Participate in the full lifecycle of HPC cluster ops from system bring-up-and-down, workload characterisation and optimisation, and rollout of new AI and Software Services.
- Design & operate a GPU orchestration layer with high availability and utilisation for AI training, inference and other scientific workloads.
- Partner with other DSO engineers to design standards, automate operations, and translate research code into performant workloads on distributed systems.
- Maintain hardware infrastructure, distributed storage, high speed networking and supporting IT infrastructure and support maintenance and upgrades.
- Degree in Computer Science & Engineering / Software Engineering / Artificial Intelligence or any other related field
- Minimum 2-year experience in IT Infrastructure or related field. More experience candidates may be considered for senior role.
- Strong proficiency in Linux environments, computer architecture, and Python / Bash scripting for tooling and automation.
- Working proficiency of Kubernetes container orchestration and infrastructure provisioning/management software (e.g. Ansible, Terraform) for fleet automation.
- Experience with NVIDIA GPUs, GitOps, Infra CI/CD, networking protocols, and other AI infrastructure technologies will be advantageous.
- Strong written and verbal communication to lead vendor and cross-functional engagements and/or performance analysis and troubleshooting initiatives.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search