Site Reliability Engineer
Indexed description
Responsibilities
Responsibilities:
- Design and maintain highly available production systems.
- Define and manage SLIs, SLOs, and error budgets.
- Automate operational tasks and eliminate manual processes.
- Develop monitoring, alerting, and observability solutions.
- Improve system performance, capacity, and resilience.
- Lead incident response and root cause analysis.
- Implement disaster recovery and continuity strategies.
- Partner with development teams to improve application reliability.
Required Skills And Experience
- 5–10+ years of engineering experience, with a strong background in Linux and Windows systems
- Expertise in Kubernetes and container platforms
- Experience working with cloud infrastructure environments
- Proficiency in scripting languages such as Python and Go
- Hands-on knowledge of Terraform and automation tools
- Familiarity with monitoring platforms and incident management practices
- Experience designing and managing CI/CD pipelines
- Kubernetes certifications
- AWS/Azure certifications
- DevOps certifications
- ITIL preferred
#DICE #Bluestone
Posted Salary Range: USD $230,000.00 - USD $250,000.00 /Yr.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search