Senior Site Reliability Engineer, DGX Cloud
Indexed description
NVIDIA is driving AI and high-performance computing forward. DGX Cloud aims to deliver a fully managed AI platform on major cloud providers, optimizing AI workloads using high-performance NVIDIA infrastructure. Work with NVIDIA's DGX Cloud team as a Senior Site Reliability Engineer to maintain high-performance DGX Cloud clusters for AI researchers and enterprise clients worldwide.
What You’ll Be Doing
- Build, implement and support operational and reliability aspects of large-scale Kubernetes clusters with focus on performance at scale, real time monitoring, logging and alerting
- Define SLOs/SLIs, monitor error budgets, and streamline reporting
- Support services before they launch through system creation consulting, developing software tools, platforms and frameworks, capacity management, and launch reviews
- Maintain services once they are live by measuring and monitoring availability, latency and overall system health
- Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds
- Scale systems sustainably through mechanisms like automation and evolve systems by pushing for changes that improve reliability and velocity
- Lead triage and root-cause analysis of high-severity incidents
- Practice balanced incident response and blameless postmortems
- Participate in on-call rotation to support production services
- BS in Computer Science or related technical field, or equivalent experience
- 10+ years of experience operating production services
- Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture
- Experience with infrastructure automation tools (e.g., Terraform, Ansible, Chef, Puppet)
- Proficiency in at least one high-level programming language (e.g., Python, Go)
- In-depth knowledge of Linux operating systems, networking fundamentals (TCP/IP), and cloud security standards
- Proficient knowledge of SRE principles, encompassing SLOs, SLIs, error budgets, and incident handling
- Experience building and operating comprehensive observability stacks (monitoring, logging, tracing) using tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc.
- Operating GPU-accelerated clusters with KubeVirt in production
- Applying generative-AI techniques to reduce operational toil
- Experience with workflow orchestration platforms such as Temporal, Cadence, Airflow, Argo Workflows, or Step Functions
- Experience operating and troubleshooting production AI inference workloads across the model-to-GPU stack, including vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, NCCL, and GPU performance analysis
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search