AI Platform Engineer
Indexed description
DAMAC AI is running NVIDIA NCP-compliant GPU infrastructure — B300/GB300 Blackwell clusters. Someone needs to build the platform layer that turns that raw compute into something data scientists and ML engineers can actually use. That's this role.
The Role
You will design and build the orchestration, scheduling, and tooling layer that sits between GPU hardware and AI workloads — Kubernetes-based GPU orchestration, MLOps enablement, platform observability, and developer experience. This is greenfield platform engineering at the hardware-to-workload boundary.
What You Will Do
- Design AI platform architecture bridging GPU hardware and workload layers
- Build Kubernetes-based GPU orchestration — scheduling, resource allocation, multi-tenancy
- Build self-service MLOps tooling for data scientists and ML engineers
- Own platform observability and performance monitoring
- Implement platform security and compliance for multi-tenant environments
- Drive automation and developer experience improvements
What You Bring
- Strong Kubernetes experience with GPU orchestration — device plugins, MIG, time-slicing
- Platform or MLOps tooling experience abstracting infrastructure for ML users
- Direct NVIDIA GPU infrastructure experience — DGX/HGX or NCP-compliant environments
- Platform architecture experience at the hardware-to-workload boundary
- Strong automation and CI/CD skills
- Comfortable building from scratch in a fast-moving environment
Dubai-based. Immediate requirement.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search