Senior Kubernetes Platform Developer – GPU & AI Infrastructure
Indexed description
Senior Kubernetes Platform Developer – GPU & AI Infrastructure
Location: Dallas, TX preferred
Work Arrangement: Hybrid, 3 days onsite / 2 days remote
Remote Flexibility: Full remote may be considered for the right candidate
Relocation: Available
Employment Type: Direct Hire
Overview
We are seeking a Senior Kubernetes Platform Developer to design and build the software powering a next-generation GPU-accelerated compute platform supporting AI, machine learning, LLM, and HPC workloads.
This is a software development role focused on Kubernetes, not a traditional DevOps, SRE, or Kubernetes administration position.
The core focus is developing Kubernetes-native software including custom operators, controllers, CRDs, APIs, scheduling capabilities, and internal platform services used to orchestrate large-scale GPU infrastructure.
The ideal candidate is a strong developer who understands Kubernetes internals and has experience building software on top of Kubernetes, not simply deploying applications or maintaining clusters.
Key Responsibilities
- Develop Kubernetes-native software using Go, Python, or similar languages.
- Build custom operators, controllers, CRDs, APIs, and platform services.
- Extend Kubernetes to support GPU-intensive AI/ML and HPC workloads.
- Develop automation for cluster provisioning, lifecycle management, scheduling, and infrastructure orchestration.
- Build GPU scheduling, allocation, workload placement, and resource-isolation capabilities.
- Integrate NVIDIA technologies including GPU Operator, device plugins, MIG, and DCGM.
- Develop internal tools and APIs for provisioning and managing GPU compute resources.
- Improve platform scalability, GPU utilization, workload performance, and reliability.
- Integrate Kubernetes with high-performance networking, storage, and bare-metal infrastructure.
- Build observability and automated remediation capabilities for distributed compute environments.
Required Qualifications
- Strong software development experience with Go, Python, or another modern programming language.
- Hands-on experience building Kubernetes operators, controllers, CRDs, APIs, or other Kubernetes-native software.
- Deep understanding of Kubernetes architecture, controllers, reconciliation, scheduling, RBAC, networking, and cluster lifecycle.
- Experience developing platforms or distributed systems built on Kubernetes.
- Experience with GPU infrastructure and NVIDIA technologies.
- Experience supporting AI/ML, LLM, HPC, or other compute-intensive workloads.
- Strong Linux and distributed systems knowledge.
- Experience with Terraform, Helm, Kustomize, Argo CD, Flux, or similar tooling.
- Ability to troubleshoot across Kubernetes, compute, networking, storage, GPUs, and applications.
Preferred Qualifications
- Experience with NVIDIA GPU clusters.
- Experience with Slurm, Volcano, kube-scheduler extensions, or custom scheduling.
- Familiarity with CUDA, NCCL, PyTorch, or TensorFlow.
- Experience with InfiniBand, RDMA, RoCE, or high-performance networking.
- Experience with bare-metal Kubernetes.
- Experience building internal developer platforms or self-service infrastructure.
- Background in AI infrastructure, HPC, cloud infrastructure, or large-scale distributed systems.
Ideal Candidate
The ideal candidate is a platform developer who builds Kubernetes-native systems.
This person should be comfortable writing operators, controllers, APIs, schedulers, and automation that extend Kubernetes and manage complex GPU infrastructure.
Candidates whose experience is primarily DevOps, CI/CD, Terraform administration, application deployment, or Kubernetes operations without substantial software development experience are unlikely to be the right fit.
Dallas-based candidates are preferred, but full remote may be considered for candidates with exceptional Kubernetes development and GPU infrastructure experience.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search