Senior Platform Engineer - Core Infrastructure
Indexed description
If you'd like to build the world's best AI cloud, join us.
- Note: This position requires presence in our San Francisco,San Jose, or Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.
What You’ll Do
- Architect, deploy, and operate Kubernetes clusters across Lambda's bare-metal datacenters.
- Build and maintain automation for cluster lifecycle management: provisioning, upgrades, and scaling.
- Similar to SRE-level coding abilities (small scripts, k8s operators/CRDs a plus)
- Own the reliability, performance, and security of Kubernetes workloads in production.
- Implement observability, logging, and alerting for clusters and critical workloads.
- Partner with product teams to design scalable, cloud-native services.
- Set the standards for resource management, networking, and RBAC across the platform.
- Lead incident response, root-cause analysis, and post-mortems for platform issues.
- Mentor junior engineers and raise the bar for platform engineering across the org.
- 5+ years in Platform, Infrastructure, or SRE roles, including running Kubernetes in production at scale.
- Deep knowledge of Kubernetes internals and day-2 operations (upgrades, scaling, troubleshooting).
- Strong with Helm, Kustomize, or similar, and GitOps-based delivery.
- Proficient with Linux systems and understanding of system-level operations.
- Proficient with infrastructure-as-code (Terraform, Pulumi, or equivalent).
- Solid grounding in networking, service meshes, and container runtimes.
- Hands-on with observability stacks (Prometheus, Grafana, OpenTelemetry).
- Strong coding skills in Go or Python for automation and tooling.
- Practical security experience: network policies, secrets management, and image scanning.
- Experience with multi-cluster, multi-cloud, or hybrid environments.
- Knowledge of GPU scheduling, HPC workloads, or ML/AI infrastructure.
- Experience with workflow orchestration / durable execution frameworks (Temporal, Cadence, or Argo Workflows).
- Exposure to cost optimization and capacity planning for large clusters.
- Contributions to CNCF or Kubernetes open-source projects.
- CKA/CKS certification.
About Lambda
- Founded in 2012, with 500+ employees, and growing fast
- Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
- We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
- Our values are publicly available: https://lambda.ai/careers
- We offer generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan that we all actually use
Compensation Range: $230K - $340K
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search