Risewave Consulting, Inc.
Linkedin · Posted 4d ago
Technical Manager - GPU Cloud & AI Infrastructure
Continue to application
Add your email once, then Caio opens the original posting.
Indexed description
Key Responsibilities
- GPU Infrastructure Architecture
- Drive end-to-end technical architecture design and review for GPU servers and AI compute clusters.
- Participate in the evaluation and selection of mainstream high-performance GPU hardware platforms and server architectures.
- Assess server topologies, CPUs, GPUs, system memory, high-speed local NVMe storage, and network interfaces.
- Design and review high-speed networking fabrics, including InfiniBand, RoCE, and high-speed Ethernet topologies.
- Evaluate NVLink, NVSwitch, RDMA, and multi-node/multi-GPU interconnect architectures.
- Work with data centre teams to define power density, liquid/air cooling, PDU distribution, network topology, and cabling specifications.
- Establish standard operating procedures for cluster deployment, benchmark testing, acceptance, and expansion.
- GPU Cloud Platform Development
- Steer the technical strategy and buildout of GPU resource pools, compute clusters, and cloud platform services.
- Deploy and maintain GPU resource scheduling and orchestration platforms built on Kubernetes and/or Slurm.
- Manage GPU drivers, parallel acceleration libraries, container runtimes, GPU Operators, and underlying software stacks.
- Architect multi-tier product offerings spanning Bare Metal, Virtual Machines, Containerized Instances, and On-Demand GPU slots.
- Implement full-card allocation, MIG slicing, GPU sharing, and multi-tenant resource isolation.
- Advance the development of user access control, quota management, automated provisioning, billing integration, and API services.
- Collaborate closely with product and business teams to bring standardized GPU Cloud products to market.
- AI Workload Enablement
- Support enterprise customers with LLM training, fine-tuning, inference deployment, and performance optimization.
- Analyze customer workload specifications, including model parameter scales, dataset size, concurrency, throughput, and latency SLAs.
- Recommend optimal GPU models, cluster node counts, network fabrics, and storage configurations tailored to client workloads.
- Support mainstream AI frameworks (PyTorch, TensorFlow) and distributed training frameworks.
- Support high-performance inference frameworks and microservices such as vLLM, Triton, TensorRT-LLM, and AI microservice architectures.
- Lead proofs-of-concept (PoCs), performance benchmarking, and technical acceptance testing.
- Platform Operations & SLA Management
- Build comprehensive monitoring and alerting systems across GPUs, compute nodes, network switches, and storage (e.g., DCGM, Prometheus, Grafana).
- Track GPU utilization, VRAM usage, power consumption, thermal profiles, ECC hardware errors, and idle resource rates.
- Implement metering logic based on GPU-hours, instance-hours, or token generation metrics.
- Define Service Level Agreements (SLAs), incident severity levels, escalation pathways, and disaster recovery plans.
- Continuously optimize platform uptime, compute efficiency, and commercial output per GPU node.
- Solutions Architecture & Pre-Sales Support
- Partner with sales and business development teams during customer technical discovery calls and solution proposals.
- Translate client AI requirements into technical architecture blueprints, Bills of Materials (BOMs), RFPs/RFQs, and execution roadmaps.
- Deliver technical presentations, platform demonstrations, PoCs, and technical onboarding for strategic clients.
- Define project technical boundaries, deliverables, SLAs, and customer acceptance criteria.
- Vendor & Project Management
- Manage technical delivery with server OEMs, data centre operators, network/storage vendors, and system integrators.
- Oversee hardware delivery, rack mounting, cabling, environment initialization, cluster commissioning, and acceptance.
- Manage project timelines, mitigate technical risks, and drive issue resolution.
- Standardize technical documentation, operational runbooks, and delivery guidelines.
- Build and lead a team of infrastructure and platform engineers as the business scales.
Job Requirements
- Bachelor’s degree or above in Computer Science, Electronic Engineering, Telecommunications, Software Engineering, or a related discipline.
- A minimum of 5 years of experience in Cloud Computing, High-Performance Computing (HPC), AI Infrastructure, or Data Centre Engineering.
- At least 2 years of hands-on experience with GPU clusters, AI cloud platforms, or large-scale compute environments.
- Strong proficiency in Linux, Docker, Kubernetes, Helm, and cloud-native container ecosystems.
- In-depth knowledge of GPU hardware architectures, parallel compute libraries, driver stacks, and AI server architectures.
- Hands-on experience with Kubernetes GPU scheduling or Slurm cluster management.
- Solid understanding of high-speed interconnect technologies, including InfiniBand, RoCE, RDMA, and NVLink.
- Practical understanding of AI workload execution spanning LLM training, fine-tuning, and inference serving.
- Strong client-facing communication, cross-functional collaboration, and vendor management skills.
- Fluent in both English and Mandarin, with the ability to conduct technical meetings and produce technical documentation in both languages.
- Based in Malaysia with willingness to travel within the Southeast Asia region in future.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search
Want help applying to roles like this?
Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search