Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote
Indexed description
As the company continues to scale its AI infrastructure platform, reliability has become a critical business function. This role sits at the center of that effort, partnering with Infrastructure, Product Engineering, and Support teams to improve uptime, strengthen observability, establish SLOs, reduce operational toil through automation, and lead incident response initiatives. The ideal candidate brings experience supporting large-scale production environments and enjoys solving complex reliability challenges while influencing engineering practices across a rapidly growing organization. This is an opportunity to gain exposure to cutting-edge AI and GPU infrastructure, take ownership of high-impact initiatives, and help shape the reliability strategy of a platform relied upon by more than one million developers.
Required Skills & Experience
- 5+ years of experience within major public cloud environment like AWS, GCP
- Strong Linux systems administration experience
- Strong networking fundamentals and troubleshooting skills
- Experience supporting containerized environments (Kubernetes preferred)
- Experience with monitoring, alerting, and observability tools
- Experience defining and managing SLIs, SLOs, and reliability metrics
- Incident response and postmortem experience
- Scripting or programming experience. Python, Go, Bash, or similar technologies
- Distributed systems and failure scenarios
- Kubernetes
- Prometheus, Grafana, or similar monitoring platforms
- Experience supporting GPU infrastructure or AI/ML platforms
- Infrastructure as Code experience (Terraform preferred)
- CI/CD pipeline experience
- 40% Linux & Kubernetes Administration
- 25% Monitoring, Observability & Incident Response
- 20% Automation & Reliability Engineering
- 15% Distributed Systems & Cloud Infrastructure
- 80% Hands-On Engineering
- 5% Management Duties
- 15% Team Collaboration
- medical, dental, and vision benefits
- Equity / Stock Options
- Remote equipment stipend
- Annual learning and development budget
- Flexible PTO
- Career Growth Within a Rapidly Scaling AI Infrastructure Company
Posted By: John Gordon
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search