AI Infrastructure Lead – GPU & AI Compute
Indexed description
About the Role:
We are looking for an experienced AI Factory Infrastructure Lead to design, build, and operate the infrastructure foundation for our AI Factory / Token Factory.
This is a senior technical role for someone who understands how to architect GPU-based AI compute platforms at scale — from individual accelerator performance through to multi-rack clusters, high-performance networking, storage, power, cooling, orchestration, and operational reliability.
The successful candidate will help turn AI infrastructure ambition into a resilient, efficient, and scalable compute platform capable of supporting demanding training and inference workloads.
What You Will Do:
Lead GPU Technology and Architecture
- Maintain deep technical understanding of current and emerging GPU architectures.
- Evaluate different generations and configurations of NVIDIA and other relevant AI accelerators.
- Assess GPU and AI accelerator platforms across memory bandwidth, tensor performance, interconnect capability, training and inference performance, power efficiency, and total cost of ownership.
- Develop GPU technology and lifecycle roadmaps.
Design Scalable AI Cluster Architecture
- Design scalable GPU clusters ranging from smaller development environments to large-scale multi-rack systems.
- Design compute, networking, storage, control, and management layers.
- Architect the compute, networking, storage, control, and management layers required for high-performance AI workloads.
Shape Data-Center Engineering Requirements
- Define rack layouts and infrastructure requirements.
- Define power requirements, cooling approach, redundancy model, UPS requirements, and network topology for high-density AI infrastructure.
- Work closely with data-center and facilities teams.
Drive Capacity, Performance, and Utilization
- Develop capacity models for training and inference.
- Translate business demand into technical capacity requirements, including GPU hours, token throughput, memory, network, and storage requirements.
- Optimize GPU utilization and workload scheduling.
- Benchmark models and inference engines across alternative hardware configurations.
Operate Reliable AI Infrastructure
- Establish monitoring and operational processes for GPU utilization, memory, temperature, power, networking, workload performance, and hardware health.
- Establish lifecycle, maintenance, firmware, and driver-management processes.
- Define availability and disaster-recovery architecture.
Optimize Cost and Commercial Efficiency
- Build cost models covering cost per GPU hour, cost per million tokens, power cost, infrastructure utilization, and total cost of ownership.
- Compare on-premise, hosted, and public-cloud economics.
What We Are Looking For:
Required Experience and Capabilities
- Strong background in large-scale compute, HPC, GPU computing, cloud infrastructure, or data-center architecture.
- Deep understanding of GPU architecture and AI workloads.
- Strong understanding of high-performance networking and distributed systems.
- Experience designing compute clusters and production infrastructure.
- Strong knowledge of Linux, containers, Kubernetes, and infrastructure automation.
- Ability to perform detailed capacity and performance engineering.
Preferred Experience
- Hands-on experience with NVIDIA H100/H200/B-series or equivalent platforms.
- Experience with DGX, HGX, SuperPOD, or comparable AI infrastructure.
- Knowledge of CUDA, NCCL, TensorRT, Triton, vLLM or related technologies.
- Experience with InfiniBand, RDMA, NVLink, NVSwitch and high-speed Ethernet.
- Experience with direct-to-chip liquid cooling and high-density AI data centers.
- Experience designing clusters at several hundred GPUs or greater.
Why Join Us:
This is an opportunity to play a defining role in building a high-performance AI infrastructure platform from the ground up.
You will work at the intersection of AI, cloud, data-center engineering, and advanced GPU computing, with direct influence over the architecture, performance, reliability, and economics of a strategic AI platform.
Interested in building AI infrastructure at scale? Apply now and join us in shaping the next generation of AI infrastructure.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search