Senior Site Reliability Engineer - Fleet
Indexed description
If you'd like to build the world's best AI cloud, join us.
- Note: This position requires presence in our San Francisco or Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.
- Build and operate monitoring and alerting for cluster health — fabric, GPU, power/thermal, and job-level signals — to detect and respond to issues proactively
- Remotely deploy and configure large-scale HPC clusters for AI workloads using automation wherever possible
- Automate cluster lifecycle: operating systems, firmware, drivers, and networking, managed as code (Ansible, Terraform) rather than by hand
- Create runbooks and automated remediations for common cluster failure modes, designed so Support and HPC Support can run them safely
- Troubleshoot and resolve cluster issues across InfiniBand/RoCE, NCCL, GPU-direct, fabric, switching, and power — working closely with on-site deployment teams
- Participate in on-call rotations and lead incident response for cluster-level problems
- Contribute to and maintain Standard Operating Procedures, and feed clear requirements back to other engineering teams on simplification, stability, and operational efficiency
- 7+ years of experience in Site Reliability Engineering, HPC Engineering, DevOps, or a similar role
- Have a strong understanding of modern AI infrastructure, from GPU architectures to hardware performance optimization
- Strong understanding of Linux-based systems in a distributed environment
- Are experienced configuring and troubleshooting InfiniBand (IB), RoCE, CLOS fabrics, 100GbE, Ethernet/switching, GPU-direct, and NCCL environments
- Solid understanding of Python and Go, with experience working with SWE teams to improve internal tooling.
- Experience with monitoring and alerting tools (e.g., Prometheus, Grafana, Clickhouse)
- Proficiency in automation and configuration management tools (e.g., Ansible, Terraform)
- Have excellent problem-solving and troubleshooting skills and an innate attention to detail
- Passion for continuous improvement and innovation
- Experience with machine learning / deep learning frameworks (PyTorch, TensorFlow) and benchmarking tools (DeepSpeed, MLPerf)
- Knowledge of containerization and orchestration technologies (e.g., Docker, Kubernetes)
- Experience building and/or operating HPC resources.
- Depth in the NVIDIA hardware and firmware ecosystem
- Experience with data center power and thermal design
- Background in chaos engineering or similar reliability testing methodologies
- Understanding of compliance frameworks (SOC 2, ISO 27001, etc.)
About Lambda
- Founded in 2012, with 500+ employees, and growing fast
- Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
- We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
- Our values are publicly available: https://lambda.ai/careers
- We offer generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan that we all actually use
Compensation Range: $240K - $356K
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search