Data Center Operations Engineer
Indexed description
Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit https://ir.bitdeer.com/
Key Responsibilities
- Responsible for the daily operation and maintenance of the Data Center infrastructure to ensure high availability and stable service operation.
- Perform installation, rack and stack, cabling, commissioning, maintenance, and troubleshooting of AI/HPC cluster infrastructure, including:
- NVIDIA B300 Cluster
- GPU Servers
- x86 Servers
- Storage Servers
- Ethernet and InfiniBand Switches
- DAC, AOC, Optical Fiber, and related cabling infrastructure
- Monitor and maintain the health status of cluster systems, including servers, GPUs, storage, networking devices, and associated infrastructure.
- Conduct hardware replacement and maintenance activities, including FRU replacement, BIOS/BMC/Firmware upgrades, and hardware diagnostics.
- Support server provisioning, operating system installation, cluster expansion, network validation, and burn-in testing.
- Troubleshoot hardware and infrastructure issues, including server failures, GPU errors, storage issues, network connectivity problems, switch failures, and cabling faults.
- Perform routine inspections, preventive maintenance, and maintain accurate operational records and maintenance logs.
- Execute incident response procedures and provide timely escalation and resolution according to operational standards.
- Prepare shift handover reports and maintain operation documents, SOPs, and incident reports.
- Work closely with engineering, network, and infrastructure teams to support new deployments and ongoing operation improvements.
- Participate in a three-shift rotation schedule, including night shifts, weekends, and holidays as required.
- Bachelor's degree or above in Computer Science, Computer Engineering, Electrical Engineering, Electronics Engineering, Information Technology, or related disciplines.
- Basic understanding of Data Center infrastructure and server hardware architecture.
- Familiarity with one or more of the following systems:
- NVIDIA B300 Cluster
- GPU Servers
- x86 Servers
- Storage Servers
- Ethernet and InfiniBand Networks
- Knowledge of server hardware components, including CPU, memory, storage, GPU, BMC/IPMI, and firmware management.
- Familiarity with network concepts, including TCP/IP, Ethernet, VLAN, Link Aggregation (LACP), and high-speed interconnect technologies such as InfiniBand or RoCE.
- Understanding of structured cabling systems, including DAC, AOC, optical fiber, MPO, and LC connectors.
- Basic Linux administration skills, including:
- System monitoring and troubleshooting
- Service management using systemctl
- Log analysis using journalctl and dmesg
- Network troubleshooting tools such as ip and ethtool
- Basic shell scripting
- Experience in Data Center operations or hardware maintenance is preferred.
- Experience supporting AI/HPC infrastructure or GPU clusters is a plus.
- Familiarity with NVIDIA AI infrastructure, including GB200 and GB300 systems, is highly desirable.
- Experience with large-scale cluster environments and high-speed networking technologies is a plus.
- Familiarity with monitoring and orchestration tools such as Slurm, Kubernetes, Prometheus, or Grafana is an advantage.
- Willingness to work in a 24x7 shift rotation schedule, including night shifts.
- Strong sense of responsibility and ownership.
- Good teamwork and communication skills.
- Ability to work under pressure and respond effectively to operational incidents.
- Detail-oriented with strong adherence to operational procedures and safety standards.
- Self-motivated with a proactive attitude toward learning and problem-solving.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search