Only W2: SRE / Production Reliability Engineer, Woonsocket, RI
Indexed description
The ideal candidate should have strong experience in SRE/DevOps, Incident Management, Observability, Kubernetes, GCP, and production monitoring, along with hands-on experience with time-series anomaly detection.
What You Ll Do
- Own reliability and performance of critical production applications.
- Act as an Incident Commander (IC) during P1/P2 production incidents.
- Lead incident response, root-cause analysis, postmortems, and reliability improvements.
- Define and manage SLIs, SLOs, and error budgets.
- Build and improve monitoring, alerting, and observability solutions.
- Tune and validate time-series anomaly detection models for production monitoring.
- Develop automation and operational tools using Python, Java, and React.
- Troubleshoot Kubernetes, cloud, batch processing, and data pipeline issues.
- Work with engineering and operations teams to improve system reliability and reduce manual work.
- Support large-scale deployments and manage production risks such as configuration drift and blast radius.
- 8+ years of experience in SRE, DevOps, Platform Engineering, or Production Engineering.
- Hands-on experience as an Incident Commander for P1/P2 incidents.
- Strong experience with time-series anomaly detection models in production observability mandatory.
- Strong production-level programming skills in:
- Python
- Java
- React
- Strong experience with SLI, SLO, and error budgets.
- Hands-on observability experience with:
- Prometheus
- Grafana
- OpenTelemetry
- At least 2 log platforms such as Loki, Splunk, or Elasticsearch
- Strong GCP experience.
- Strong Kubernetes operational experience.
- Experience with Rancher K3s.
- Experience troubleshooting Apache Airflow and Tidal workflows/batch jobs.
- Experience with production-scale distributed systems and on-call support.
- Production Readiness Reviews / service launch experience.
- BigQuery and PostgreSQL.
- Chaos Engineering / fault injection.
- TIC/Technical Incident Commander certification.
- Healthcare, pharmacy, retail, or other high-availability environments.
- LLM/GenAI for incident management, alert summarization, or runbook recommendations.
- Kafka.
- Istio / Envoy.
- Terraform / Ansible.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search