Data Center Site Manager / Supervisor
Indexed description
Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit https://ir.bitdeer.com/
Key Responsibilities
- Lead and manage the daily operations of the Data Center site, ensuring the availability, reliability, and operational excellence of all infrastructure and systems.
- Supervise and manage a team of Operations Engineers, including manpower planning, shift scheduling, task assignment, performance management, coaching, and professional development.
- Ensure 24x7 operational coverage and maintain adequate staffing to support business and customer requirements.
- Act as the primary escalation point for operational incidents and coordinate cross-functional teams to drive timely issue resolution and root cause analysis.
- Oversee the operation, maintenance, and troubleshooting of AI/HPC infrastructure, including:
- NVIDIA B300 Clusters
- GPU Servers
- x86 Servers
- Storage Servers
- Ethernet and InfiniBand Switches
- Optical fiber, DAC, AOC, and associated cabling systems
- Establish, maintain, and continuously improve operational procedures, Standard Operating Procedures (SOPs), Emergency Operating Procedures (EOPs), and preventive maintenance programs.
- Monitor site health, operational KPIs, incident trends, and infrastructure performance to ensure service quality and operational efficiency.
- Coordinate hardware installation, rack and stack activities, system commissioning, infrastructure expansion, and lifecycle management.
- Review and approve maintenance activities, change requests, incident reports, and shift handover records.
- Ensure compliance with company policies, operational standards, safety requirements, and security procedures within the Data Center.
- Collaborate with engineering, network, facilities, and vendor teams to support new deployments and operational improvement initiatives.
- Participate in on-call rotation and provide hands-on operational support when necessary, including covering shift duties during manpower shortages, emergencies, or critical incidents.
- Drive a culture of operational excellence, teamwork, accountability, and continuous improvement within the site operation team.
- Bachelor's degree or above in Computer Science, Computer Engineering, Electrical Engineering, Electronics Engineering, Information Technology, or related disciplines.
- Minimum 5 years of experience in Data Center operations, IT infrastructure, or HPC/AI infrastructure management.
- Minimum 2 years of experience in team leadership or people management.
- Proven experience managing 24x7 shift operations in a mission-critical environment is preferred.
- Experience with large-scale AI or HPC clusters is highly desirable.
- Strong knowledge of Data Center operations and infrastructure management, including:
- NVIDIA B300 Clusters
- GPU Servers
- x86 Servers
- Storage Systems
- Ethernet Networking
- InfiniBand Networking
- Familiarity with NVIDIA GPU architecture, NVLink, NVSwitch, and AI cluster deployment concepts.
- Experience with server hardware troubleshooting, firmware management, and hardware lifecycle management.
- Good understanding of structured cabling systems, including optical fiber, MPO/LC connectors, DAC, and AOC cabling.
- Familiarity with infrastructure monitoring and management tools.
- Strong Linux administration and troubleshooting skills, including:
- System and service management
- Hardware and performance diagnostics
- Log analysis and incident investigation
- Network troubleshooting
- Basic scripting and automation
- Demonstrated ability to lead, motivate, and develop a team of Operations Engineers.
- Experience in:
- Shift scheduling and workforce planning
- Incident and escalation management
- Performance management and coaching
- SOP/EOP development and operational governance
- Vendor and stakeholder coordination
- Strong decision-making skills with the ability to manage priorities in a fast-paced and mission-critical environment.
- Willingness to provide hands-on operational support and participate in on-call duties when required.
- Strong ownership mindset and accountability for site operations.
- Excellent communication and interpersonal skills.
- Ability to remain calm and make sound decisions during critical incidents.
- Highly organized, detail-oriented, and committed to operational excellence and continuous improvement.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search