Head of Compute & Inference Platform
Indexed description
This role is responsible for building the platform that efficiently shares GPU resources, serves AI models at scale, and delivers high-performance, multi-tenant AI infrastructure to customers. You'll own everything from workload scheduling and GPU allocation to model serving, runtime optimization, APIs, and platform performance.
What You'll Do
Compute & Inference Platform Strategy
- Define the technical vision and roadmap for Nava's Compute & Inference Platform.
- Own the architecture and evolution of GPU scheduling, workload orchestration, and inference serving infrastructure.
- Build a highly scalable, secure, and multi-tenant AI platform capable of serving enterprise workloads.
- Drive platform innovation while balancing performance, reliability, cost, and customer experience.
- Design and optimize GPU scheduling, allocation, and capacity management across large-scale GPU clusters.
- Develop intelligent scheduling strategies to maximize GPU utilization while maintaining fairness and workload isolation.
- Own resource quotas, workload prioritization, tenancy management, and capacity planning.
- Continuously improve infrastructure efficiency and cost optimization.
- Own the end-to-end model serving stack, including deployment, scaling, and lifecycle management.
- Lead engineering for model runtimes, inference frameworks, serving infrastructure, and API gateways.
- Ensure rapid onboarding and deployment of new AI models while maintaining platform stability and performance.
- Optimize inference latency, throughput, and infrastructure utilization.
- Own customer-facing APIs, SDKs, and platform interfaces that enable seamless deployment and management of AI workloads.
- Improve developer experience through automation, self-service capabilities, and platform tooling.
- Work closely with Product teams to define platform capabilities and customer-facing features.
- Establish platform performance benchmarks, service level objectives (SLOs), and cost optimization targets.
- Continuously improve GPU utilization, inference efficiency, scheduling algorithms, and workload performance.
- Drive benchmarking, performance testing, and capacity planning initiatives.
- Build observability and telemetry to measure platform health and customer experience.
- Collaborate closely with GPU Cluster Engineering, Platform Reliability, AI Infrastructure Security, Networking, Product, and Customer Success teams.
- Align infrastructure capabilities with product strategy and customer requirements.
- Act as the technical leader for platform architecture and major engineering decisions.
- Build and lead a high-performing Compute & Inference engineering organization.
- Mentor engineering managers and senior engineers.
- Foster a culture of technical excellence, ownership, innovation, and operational discipline.
- Infrastructure Efficiency: Achieve and sustain ≥85% average GPU utilization across production clusters.
- Inference Performance: Deliver sub-50ms p99 latency for high-volume models, with ≥95% SLA adherence for throughput targets.
- Developer Velocity: Reduce time-to-deploy new models from days to minutes, supporting ≥200 model deployments/month.
- Reliability & Isolation: Maintain ≥99.95% platform uptime and enforce strict performance isolation across tenants (≤5% variance under load).
- Customer Experience: Achieve ≥4.5/5 NPS on platform usability and ≥99.9% API availability.
- Cost Efficiency: Decrease cost per inference by ≥20% YoY while scaling throughput by ≥2x.
- Time-to-Value: Reduce time-to-production for new models by ≥50% over 12 months.
- 12+ years of experience building large-scale distributed systems, cloud platforms, AI infrastructure, or compute platforms.
- Proven experience leading platform engineering teams responsible for production-scale infrastructure.
- Deep expertise in:
- GPU scheduling and resource management
- Kubernetes and container orchestration
- Distributed systems and cloud-native platforms
- AI model serving and inference architectures
- Multi-tenant platform design
- API platforms and developer tooling
- Performance engineering and capacity planning
- Strong understanding of AI infrastructure technologies, including model serving frameworks and orchestration platforms.
- Excellent architectural thinking with the ability to balance scalability, performance, security, and operational simplicity.
- Exceptional leadership, stakeholder management, and communication skills.
- Experience with NVIDIA GPU ecosystems, CUDA, Triton Inference Server, vLLM, Ray Serve, KServe, or similar inference technologies.
- Familiarity with LLM serving, distributed inference, model optimization, and quantization techniques.
- Experience building AI cloud platforms, GPU-as-a-Service offerings, or hyperscale infrastructure.
- Exposure to HPC environments and large-scale enterprise AI deployments.
- Lead the engineering vision for one of the world's most advanced AI inference platforms.
- Solve some of the most challenging problems in GPU scheduling, distributed inference, and AI infrastructure.
- Work alongside world-class engineers building the future of enterprise AI.
- Shape the platform that powers next-generation AI applications at global scale.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search