Senior Site Reliability Engineer
Indexed description
- SF Bay Area / Remote (US)
This is a hands-on, close-to-the-metal role for a first-principles Linux engineer. You'll be the final escalation for the hardest GPU, networking, and kernel-level failures, sometimes debugging directly with NVIDIA. It fits someone who thrives on low-level problems in a fast, less-structured environment. If you want a narrow, well-bounded ops role, this isn't it.
What You'll Own
- Take end-to-end ownership of production GPU clusters for training and inference across AWS and OCI, keeping them highly available and performant.
- Join critical re-architecture sessions to redesign systems for higher efficiency and scale.
- Tune Linux performance deeply, at the OS and kernel level.
- Build automation in Python, Go, or Bash to manage, monitor, and self-heal infrastructure without heavy toil.
- Serve as the final escalation for the hardest GPU, networking (InfiniBand/RDMA), and system failures, working with vendors like NVIDIA.
- Help achieve and maintain security certifications (SOC 2 Type 1 & 2, ISO) with strong infrastructure security practices.
- Days 1–30 — Immerse & Diagnose: Learn the current clusters across on-prem, AWS, and OCI, and where reliability and performance hurt most.
- Days 30–60 — Ship & Validate: Take ownership of a production cluster and ship automation or tuning that measurably improves availability or performance.
- Days 60–90 — Scale & Systemize: Contribute to the next-gen re-architecture and harden security and compliance practices.
- 5+ years as an SRE, production, or infrastructure engineer in a fast-paced, large-scale environment.
- Deep, hands-on Linux expertise, containerized systems, and low-level performance debugging.
- Working experience with Terraform, Airflow, and Ray.
- Strong experience with AWS or OCI.
- Practical experience with high-performance networking (InfiniBand, RDMA, or RoCE).
- Working knowledge of security best practices and compliance frameworks like SOC 2 and ISO.
- Comfort in a less-structured, fast-paced environment.
- Deep expertise with GPU tooling for NVIDIA and AMD (DCGM, ROCm).
- Experience managing large-scale GPU clusters for AI/ML training or inference.
- Familiarity with Kubernetes or orchestration frameworks like Ray.
- Deep expertise in data pipelines and infrastructure.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search