HPC Infrastructure Engineer
Indexed description
A seed-stage neocloud company is hiring an HPC Infrastructure Engineer to build and operate the foundation of its GPU and TPU cloud. You'll work across bare-metal systems, high-speed networking, cluster scheduling, storage, automation, and reliability, turning a large fleet of accelerators into a high-performance, dependable computing platform for customers. This is a hands-on role for someone who loves working across hardware and software, diagnosing tough performance problems, and building infrastructure from the ground up.
What You'll Do
- Design, deploy, and operate production TPU clusters
- Provision and manage Linux-based bare-metal servers at scale
- Build automated workflows for server installation, configuration, upgrades, and recovery
- Deploy and operate cluster schedulers such as Slurm or Kubernetes
- Integrate high-performance networking using InfiniBand or RoCE/RDMA
- Build and operate high-throughput storage for distributed AI workloads
- Monitor cluster health, accelerator utilization, network performance, storage performance, and job reliability
- Diagnose failures across GPUs, servers, firmware, networks, storage, schedulers, and customer workloads
- Improve cluster utilization, training performance, fault tolerance, and recovery time
- Establish production practices for change management, incident response, capacity management, and operational readiness
- Partner with data-center, network, platform, security, and customer-facing teams to launch new capacity
- Evaluate and manage infrastructure vendors, hardware suppliers, and technical partners
- Participate in an on-call rotation as the production platform grows
What We're Looking For
- Experience building or operating production HPC, supercomputing, or large-scale bare-metal infrastructure
- Strong Linux systems-engineering and debugging skills
- Experience automating physical server provisioning and configuration
- Experience with Slurm, Kubernetes, or another distributed workload scheduler
- Understanding of high-performance networking, distributed storage, and accelerator-based computing
- Experience with monitoring, incident response, performance analysis, and production reliability
- Proficiency in Python, Go, Bash, or another infrastructure automation language
- Ability to diagnose problems across layers rather than treating compute, networking, storage, and software as separate systems
- Comfort working directly with hardware, vendors, data-center teams, and customers in an early-stage environment
Especially Valuable
- NVIDIA GPU clusters, TPU systems, CUDA, NCCL, or collective communications
- InfiniBand, RoCEv2, RDMA, GPUDirect, or high-performance Ethernet
- Slurm administration and scheduling optimization
- Kubernetes for accelerator workloads
- PXE, Redfish, IPMI, BMCs, and bare-metal lifecycle management
- Ansible, Terraform, or other infrastructure-as-code systems
- Lustre, Spectrum Scale/GPFS, Ceph, BeeGFS, or other distributed storage systems
- Prometheus, Grafana, OpenTelemetry, or infrastructure observability tooling
- GPU health monitoring, firmware management, and hardware failure diagnosis
- Distributed AI training and inference workloads
- Performance benchmarking and optimization across compute, network, and storage
- Experience at a neocloud, hyperscaler, AI lab, national laboratory, or HPC center
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search