Back to search
TalentDome Staffing Linkedin · Posted 1mo ago

Site Reliability Engineer

Seattle, Washington, United States

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Senior Site Reliability Engineer (SRE)

Location: Seattle, hybrid - 2 times a week in the office

Job Type: Full-time, direct hire

Industry: High-Growth Technology / SaaS


About the Role

We are seeking a highly skilled Senior Site Reliability Engineer to drive the reliability, scalability, and performance of our client's production systems. The ideal candidate combines deep software engineering ability with system-level expertise, applying core SRE principles to reduce toil, minimize downtime, and build self-healing infrastructure across complex, high-scale environments.


Key Responsibilities

  • Reliability Engineering: Define and drive adoption of SLIs, SLOs, and error budgets across services, using them to guide engineering priorities and release decisions.
  • Automation & Toil Reduction: Build tools and automation (Python, Go, Bash) to eliminate manual operational work and enable self-service capabilities for engineering teams.
  • Infrastructure as Code (IaC): Design and maintain scalable, resilient infrastructure on AWS using Terraform, CloudFormation, or Pulumi, ensuring consistency and repeatability.
  • Observability: Architect monitoring, logging, tracing, and alerting systems (Prometheus, Grafana, Datadog, Splunk, OpenTelemetry) that give clear, actionable signals into system health.
  • Incident Management: Act as an incident commander during major outages, lead blameless postmortems, and drive systemic fixes to prevent recurrence.
  • Capacity Planning & Performance: Forecast growth, execute load/chaos testing, and tune systems proactively to stay ahead of scaling bottlenecks.
  • CI/CD & Deployment Safety: Partner with engineering teams to build safe, progressive delivery pipelines (canary, blue/green, feature flags) using tools like ArgoCD, Jenkins, or GitLab CI.
  • Security & Compliance: Embed security best practices into infrastructure and deployment pipelines, including access control, network segmentation, and vulnerability management.
  • On-Call Leadership: Participate in and help evolve on-call rotations, escalation policies, and runbooks to reduce alert fatigue and improve response times.
  • Mentorship & Culture: Champion SRE best practices across the organization, mentor engineers on reliability thinking, and influence upstream architecture decisions.


Qualifications & Experience

  • 6+ years of hands-on experience in Site Reliability Engineering, DevOps, or Backend/Systems Engineering with a track record of owning production reliability at scale.
  • Strong Software Engineering Background: Proficiency in Python, Go, or similar languages—focused on building maintainable services and tooling, not just basic scripting.
  • AWS Expertise: Deep technical knowledge of AWS core services (EC2, Lambda, RDS, VPC, IAM, S3, Terraform/CloudFormation) alongside cost and performance optimization.
  • Linux Systems Mastery: Demonstrated proficiency in performance tuning, kernel/network troubleshooting, and security hardening.
  • Container Orchestration: Hands-on experience operating Kubernetes/EKS in production environments at scale.
  • SRE Frameworks: Proven experience defining and operationalizing SLIs, SLOs, and error budgets.
  • Distributed Systems: Strong understanding of consistency, fault tolerance, failover strategies, and graceful degradation.
  • Observability & Incident Command: Experience building full-stack observability pipelines and leading major incident response efforts.


Nice to Have

  • Experience with multi-cloud environments (AWS, GCP, Azure).
  • Chaos engineering experience (Gremlin, Chaos Mesh, or custom fault-injection tooling).
  • Active AWS Certifications (DevOps Engineer Professional, Solutions Architect).
  • Familiarity with compliance frameworks (SOC2, ISO 27001, HIPAA).
  • Experience with service mesh technologies (Istio, Linkerd).
  • Background in internal platform/developer experience (DevEx) teams.


Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search