Senior Solutions Architect, GPU Cloud GenAI – Infrastructure
Indexed description
The work location for this role is in Mumbai.
What You Will Be Doing
- Design and architect scalable IaaS, PaaS, and SaaS layers for large-scale GPU cluster environments (32+ HGX/DGX nodes), spanning compute, networking, and storage orchestration.
- Build multi-tenant GPU cloud platforms with production-grade APIs, control planes, and platform services that abstract infrastructure complexity for end users and application teams.
- Develop cluster orchestration pipelines using Kubernetes (GPU operators, device plugins, multi-tenancy) and Slurm, optimizing for performance, reliability, and resource efficiency at scale.
- Define and implement best practices for GPU resource scheduling, isolation, quota management, and observability, ensuring secure multi-tenant isolation and compliance.
- Advise customers on deploying and scaling generative AI workloads (LLMs, MLLMs, RAG pipelines) on your infrastructure platforms, translating AI requirements into infrastructure specifications.
- Engage with C-level executives and infrastructure teams to understand requirements, deploy GPU clusters across on-premises and hybrid cloud environments, and drive platform adoption.
- Collaborate with NVIDIA engineering teams to resolve deep infrastructure bugs, provide feedback on platform capabilities, and influence product roadmap decisions.
- Partner with customer infrastructure teams to tune, scale, and optimize GPU clusters for cost efficiency, throughput, and AI workload performance.
- 5+ years of hands-on infrastructure or platform engineering experience, with demonstrated expertise designing and operating large-scale GPU clusters (100+ nodes).
- Deep expertise building IaaS, PaaS, and SaaS platform layers—architecting and developing infrastructure foundations, not consuming cloud services.
- Proficiency in Kubernetes (GPU operator, device plugins, multi-tenancy) and Slurm for HPC and AI workloads.
- Hands-on experience with infrastructure-as-code (Terraform, Helm, Ansible), CI/CD pipelines, and observability stacks (Prometheus, Grafana, DCGC).
- Strong coding ability in Python and/or Go/C++, building platform tooling and automation from scratch.
- Experience with cloud-native networking (InfiniBand, RoCE, RDMA) and distributed storage solutions for GPU environments.
- Excellent communication skills, credibly engaging both infrastructure engineers and C-level stakeholders on complex technical and strategic topics.
- Bachelor's degree in Computer Science, Computer Engineering, or equivalent experience.
- Working knowledge of LLM, MLLM, and RAG frameworks and how they map to infrastructure requirements.
- Hands-on experience with model serving frameworks (Triton Inference Server, vLLM, TensorRT-LLM) and inference optimization techniques.
- Proven track record optimizing infrastructure for cost efficiency, throughput, and resource utilization in multi-tenant production environments.
- Deep understanding of distributed training concepts (data parallelism, model parallelism, pipeline parallelism) from an infrastructure perspective.
- Experience deploying and managing GPU clusters in cloud environments (AWS, Azure, GCP) and on-premises infrastructure at enterprise scale
, , JR2019537
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search