Back to search
Arcadia Linkedin · Posted yesterday

HPC Infrastructure Engineer

San Francisco, California, United States

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

A seed-stage neocloud company is hiring an HPC Infrastructure Engineer to build and operate the foundation of its GPU and TPU cloud. You'll work across bare-metal systems, high-speed networking, cluster scheduling, storage, automation, and reliability, turning a large fleet of accelerators into a high-performance, dependable computing platform for customers. This is a hands-on role for someone who loves working across hardware and software, diagnosing tough performance problems, and building infrastructure from the ground up.


What You'll Do

  • Design, deploy, and operate production TPU clusters
  • Provision and manage Linux-based bare-metal servers at scale
  • Build automated workflows for server installation, configuration, upgrades, and recovery
  • Deploy and operate cluster schedulers such as Slurm or Kubernetes
  • Integrate high-performance networking using InfiniBand or RoCE/RDMA
  • Build and operate high-throughput storage for distributed AI workloads
  • Monitor cluster health, accelerator utilization, network performance, storage performance, and job reliability
  • Diagnose failures across GPUs, servers, firmware, networks, storage, schedulers, and customer workloads
  • Improve cluster utilization, training performance, fault tolerance, and recovery time
  • Establish production practices for change management, incident response, capacity management, and operational readiness
  • Partner with data-center, network, platform, security, and customer-facing teams to launch new capacity
  • Evaluate and manage infrastructure vendors, hardware suppliers, and technical partners
  • Participate in an on-call rotation as the production platform grows


What We're Looking For

  • Experience building or operating production HPC, supercomputing, or large-scale bare-metal infrastructure
  • Strong Linux systems-engineering and debugging skills
  • Experience automating physical server provisioning and configuration
  • Experience with Slurm, Kubernetes, or another distributed workload scheduler
  • Understanding of high-performance networking, distributed storage, and accelerator-based computing
  • Experience with monitoring, incident response, performance analysis, and production reliability
  • Proficiency in Python, Go, Bash, or another infrastructure automation language
  • Ability to diagnose problems across layers rather than treating compute, networking, storage, and software as separate systems
  • Comfort working directly with hardware, vendors, data-center teams, and customers in an early-stage environment


Especially Valuable

  • NVIDIA GPU clusters, TPU systems, CUDA, NCCL, or collective communications
  • InfiniBand, RoCEv2, RDMA, GPUDirect, or high-performance Ethernet
  • Slurm administration and scheduling optimization
  • Kubernetes for accelerator workloads
  • PXE, Redfish, IPMI, BMCs, and bare-metal lifecycle management
  • Ansible, Terraform, or other infrastructure-as-code systems
  • Lustre, Spectrum Scale/GPFS, Ceph, BeeGFS, or other distributed storage systems
  • Prometheus, Grafana, OpenTelemetry, or infrastructure observability tooling
  • GPU health monitoring, firmware management, and hardware failure diagnosis
  • Distributed AI training and inference workloads
  • Performance benchmarking and optimization across compute, network, and storage
  • Experience at a neocloud, hyperscaler, AI lab, national laboratory, or HPC center
Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search