Staffing Spot, Inc.
Linkedin · Posted 8d ago
Site Reliability Engineer (SRE)
Continue to application
Add your email once, then Caio opens the original posting.
Indexed description
Key Responsibilities
- Design, build, and operate highly available and scalable production systems.
- Define and implement SRE practices, standards, and reliability engineering strategies.
- Establish and manage SLIs, SLOs, SLAs, and error budgets for critical services.
- Own production reliability, availability, performance, and capacity of business-critical applications.
- Develop and maintain highly automated infrastructure and operational processes.
- Build and manage CI/CD pipelines for reliable and repeatable software delivery.
- Design, deploy, and manage containerized workloads using Docker and Kubernetes.
- Implement Infrastructure as Code using Terraform, CloudFormation, or similar technologies.
- Develop automation using Python, Go, Bash, or similar scripting/programming languages.
- Implement comprehensive monitoring, logging, alerting, and observability solutions.
- Troubleshoot complex production issues across applications, infrastructure, networking, databases, and cloud services.
- Lead incident response, root-cause analysis, and post-incident reviews.
- Identify recurring operational problems and eliminate them through automation and engineering solutions.
- Perform capacity planning, performance optimization, and scalability assessments.
- Implement disaster recovery, backup, business continuity, and high-availability strategies.
- Improve deployment reliability through automation, progressive delivery, rollback strategies, and release engineering.
- Partner with development teams to improve application reliability, resilience, and observability.
- Participate in on-call rotations and provide technical leadership during critical production incidents.
- Establish operational runbooks, documentation, and troubleshooting procedures.
- Mentor engineers and drive adoption of SRE and DevOps best practices.
- Evaluate new technologies and identify opportunities to improve engineering productivity and platform reliability.
- 10+ years of experience in SRE, DevOps, Cloud Engineering, Infrastructure Engineering, or Software Engineering.
- Strong experience managing large-scale production environments.
- Strong programming/scripting experience with Python, Go, Java, Bash, or similar languages.
- Hands-on expertise with AWS, Azure, or Google Cloud Platform (GCP).
- Strong experience with Kubernetes and Docker.
- Extensive experience with Infrastructure as Code, particularly Terraform.
- Strong understanding of Linux/Unix systems administration.
- Experience designing and managing CI/CD pipelines.
- Strong understanding of networking concepts including TCP/IP, DNS, HTTP/HTTPS, TLS, load balancing, and networking in cloud environments.
- Experience with microservices and distributed systems.
- Strong knowledge of monitoring and observability concepts.
- Experience with tools such as Prometheus, Grafana, ELK/Elastic Stack, Datadog, Splunk, or similar platforms.
- Strong experience with incident management, troubleshooting, RCA, and problem management.
- Understanding of SLOs, SLIs, SLAs, error budgets, and reliability metrics.
- Experience with production performance and capacity management.
- Strong understanding of security, access management, secrets management, and cloud security best practices.
- Excellent problem-solving, communication, and technical leadership skills.
- Experience designing highly available distributed systems at scale.
- Experience with AWS services such as EC2, EKS, S3,
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search
Want help applying to roles like this?
Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search