Principal SRE / Hybrid / Tempe
Indexed description
This is not a Principal role where you are removed from the technology. The team is looking for someone who can remain highly hands-on while also serving as a technical resource and mentor to the rest of the group. Kubernetes is at the center of the environment, with workloads running across EKS and AKS, and future initiatives may involve deeper cluster architecture and design. Today, the focus is heavily around scaling clusters, Kubernetes networking, improving existing architecture, automation, and reliability. The manager strongly values curiosity, attitude, and a willingness to learn over simply checking every technology box, making this a great opportunity for someone who wants continued technical growth while having real influence over the direction of the platform.
Required Skills & Experience
- 5+ years of experience within Site Reliability Engineering, DevOps, Platform Engineering, Cloud Engineering, or a similar infrastructure-focused role
- Strong hands-on Kubernetes experience
- Experience supporting and scaling production Kubernetes clusters
- Strong understanding of Kubernetes networking and architecture
- Experience with AWS and/or Azure cloud environments
- Infrastructure as Code experience, preferably Terraform
- Experience supporting highly available, production cloud environments
- Strong troubleshooting and incident response experience
- Experience with Linux-based environments
- Experience working across both SRE and DevOps responsibilities
- Ability to mentor engineers and provide technical guidance at a Principal level
- Strong communication skills and ability to work across engineering teams
- Curiosity and willingness to learn new technologies and cloud platforms
- Experience with both AWS EKS and Azure AKS
- Kubernetes cluster architecture or cluster design experience
- Terraform experience in large-scale environments
- Python, Go, or another programming/scripting language
- Experience building internal tooling or infrastructure automation
- Experience with observability, monitoring, SLOs/SLIs, and production reliability
- Experience improving Kubernetes scalability, networking, and performance
- Exposure to GCP
- Experience using AI-assisted engineering tools in a thoughtful and cost-conscious way
- Previous technical mentorship or leadership experience
- 40% Kubernetes / Container Platform Engineering
- 25% AWS & Azure Cloud Infrastructure
- 20% Terraform / Infrastructure Automation
- 15% Reliability, Monitoring & Production Engineering
- 50% Hands-On Engineering
- 25% Architecture / Technical Direction
- 25% Mentoring & Team Collaboration
- Medical, Dental, and Vision Insurance
- Vacation Time
- Paid Holidays
- 401(k)
- Bonus Eligibility
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search