Back to search
True Corporation Linkedin · Posted 16d ago

AI Infrastructure Lead – GPU & AI Compute

Bangkok

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

About the Role:

We are looking for an experienced AI Factory Infrastructure Lead to design, build, and operate the infrastructure foundation for our AI Factory / Token Factory.

This is a senior technical role for someone who understands how to architect GPU-based AI compute platforms at scale — from individual accelerator performance through to multi-rack clusters, high-performance networking, storage, power, cooling, orchestration, and operational reliability.


The successful candidate will help turn AI infrastructure ambition into a resilient, efficient, and scalable compute platform capable of supporting demanding training and inference workloads.


What You Will Do:

Lead GPU Technology and Architecture

  • Maintain deep technical understanding of current and emerging GPU architectures.
  • Evaluate different generations and configurations of NVIDIA and other relevant AI accelerators.
  • Assess GPU and AI accelerator platforms across memory bandwidth, tensor performance, interconnect capability, training and inference performance, power efficiency, and total cost of ownership.
  • Develop GPU technology and lifecycle roadmaps.

Design Scalable AI Cluster Architecture

  • Design scalable GPU clusters ranging from smaller development environments to large-scale multi-rack systems.
  • Design compute, networking, storage, control, and management layers.
  • Architect the compute, networking, storage, control, and management layers required for high-performance AI workloads.

Shape Data-Center Engineering Requirements

  • Define rack layouts and infrastructure requirements.
  • Define power requirements, cooling approach, redundancy model, UPS requirements, and network topology for high-density AI infrastructure.
  • Work closely with data-center and facilities teams.

Drive Capacity, Performance, and Utilization

  • Develop capacity models for training and inference.
  • Translate business demand into technical capacity requirements, including GPU hours, token throughput, memory, network, and storage requirements.
  • Optimize GPU utilization and workload scheduling.
  • Benchmark models and inference engines across alternative hardware configurations.

Operate Reliable AI Infrastructure

  • Establish monitoring and operational processes for GPU utilization, memory, temperature, power, networking, workload performance, and hardware health.
  • Establish lifecycle, maintenance, firmware, and driver-management processes.
  • Define availability and disaster-recovery architecture.

Optimize Cost and Commercial Efficiency

  • Build cost models covering cost per GPU hour, cost per million tokens, power cost, infrastructure utilization, and total cost of ownership.
  • Compare on-premise, hosted, and public-cloud economics.


What We Are Looking For:

Required Experience and Capabilities

  • Strong background in large-scale compute, HPC, GPU computing, cloud infrastructure, or data-center architecture.
  • Deep understanding of GPU architecture and AI workloads.
  • Strong understanding of high-performance networking and distributed systems.
  • Experience designing compute clusters and production infrastructure.
  • Strong knowledge of Linux, containers, Kubernetes, and infrastructure automation.
  • Ability to perform detailed capacity and performance engineering.

Preferred Experience

  • Hands-on experience with NVIDIA H100/H200/B-series or equivalent platforms.
  • Experience with DGX, HGX, SuperPOD, or comparable AI infrastructure.
  • Knowledge of CUDA, NCCL, TensorRT, Triton, vLLM or related technologies.
  • Experience with InfiniBand, RDMA, NVLink, NVSwitch and high-speed Ethernet.
  • Experience with direct-to-chip liquid cooling and high-density AI data centers.
  • Experience designing clusters at several hundred GPUs or greater.


Why Join Us:

This is an opportunity to play a defining role in building a high-performance AI infrastructure platform from the ground up.


You will work at the intersection of AI, cloud, data-center engineering, and advanced GPU computing, with direct influence over the architecture, performance, reliability, and economics of a strategic AI platform.


Interested in building AI infrastructure at scale? Apply now and join us in shaping the next generation of AI infrastructure.

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search