Site Reliability Engineer
Indexed description
Senior Site Reliability Engineer (SRE)
Location: Seattle, hybrid - 2 times a week in the office
Job Type: Full-time, direct hire
Industry: High-Growth Technology / SaaS
About the Role
We are seeking a highly skilled Senior Site Reliability Engineer to drive the reliability, scalability, and performance of our client's production systems. The ideal candidate combines deep software engineering ability with system-level expertise, applying core SRE principles to reduce toil, minimize downtime, and build self-healing infrastructure across complex, high-scale environments.
Key Responsibilities
- Reliability Engineering: Define and drive adoption of SLIs, SLOs, and error budgets across services, using them to guide engineering priorities and release decisions.
- Automation & Toil Reduction: Build tools and automation (Python, Go, Bash) to eliminate manual operational work and enable self-service capabilities for engineering teams.
- Infrastructure as Code (IaC): Design and maintain scalable, resilient infrastructure on AWS using Terraform, CloudFormation, or Pulumi, ensuring consistency and repeatability.
- Observability: Architect monitoring, logging, tracing, and alerting systems (Prometheus, Grafana, Datadog, Splunk, OpenTelemetry) that give clear, actionable signals into system health.
- Incident Management: Act as an incident commander during major outages, lead blameless postmortems, and drive systemic fixes to prevent recurrence.
- Capacity Planning & Performance: Forecast growth, execute load/chaos testing, and tune systems proactively to stay ahead of scaling bottlenecks.
- CI/CD & Deployment Safety: Partner with engineering teams to build safe, progressive delivery pipelines (canary, blue/green, feature flags) using tools like ArgoCD, Jenkins, or GitLab CI.
- Security & Compliance: Embed security best practices into infrastructure and deployment pipelines, including access control, network segmentation, and vulnerability management.
- On-Call Leadership: Participate in and help evolve on-call rotations, escalation policies, and runbooks to reduce alert fatigue and improve response times.
- Mentorship & Culture: Champion SRE best practices across the organization, mentor engineers on reliability thinking, and influence upstream architecture decisions.
Qualifications & Experience
- 6+ years of hands-on experience in Site Reliability Engineering, DevOps, or Backend/Systems Engineering with a track record of owning production reliability at scale.
- Strong Software Engineering Background: Proficiency in Python, Go, or similar languages—focused on building maintainable services and tooling, not just basic scripting.
- AWS Expertise: Deep technical knowledge of AWS core services (EC2, Lambda, RDS, VPC, IAM, S3, Terraform/CloudFormation) alongside cost and performance optimization.
- Linux Systems Mastery: Demonstrated proficiency in performance tuning, kernel/network troubleshooting, and security hardening.
- Container Orchestration: Hands-on experience operating Kubernetes/EKS in production environments at scale.
- SRE Frameworks: Proven experience defining and operationalizing SLIs, SLOs, and error budgets.
- Distributed Systems: Strong understanding of consistency, fault tolerance, failover strategies, and graceful degradation.
- Observability & Incident Command: Experience building full-stack observability pipelines and leading major incident response efforts.
Nice to Have
- Experience with multi-cloud environments (AWS, GCP, Azure).
- Chaos engineering experience (Gremlin, Chaos Mesh, or custom fault-injection tooling).
- Active AWS Certifications (DevOps Engineer Professional, Solutions Architect).
- Familiarity with compliance frameworks (SOC2, ISO 27001, HIPAA).
- Experience with service mesh technologies (Istio, Linkerd).
- Background in internal platform/developer experience (DevEx) teams.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search