Back to search
Tror - AI for everyone Linkedin · Posted 6d ago

Only W2: SRE / Production Reliability Engineer, Woonsocket, RI

Woonsocket

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Role: Senior SRE / Production Reliability Engineer

Experience: 8+ Years

Location: Woonsocket, RI

Job Summary

We are looking for a Senior SRE / Production Reliability Engineer to improve the reliability, performance, and availability of critical production systems.

The ideal candidate should have strong experience in SRE/DevOps, Incident Management, Observability, Kubernetes, GCP, and production monitoring, along with hands-on experience with time-series anomaly detection.

What You Ll Do

  • Own reliability and performance of critical production applications.
  • Act as an Incident Commander (IC) during P1/P2 production incidents.
  • Lead incident response, root-cause analysis, postmortems, and reliability improvements.
  • Define and manage SLIs, SLOs, and error budgets.
  • Build and improve monitoring, alerting, and observability solutions.
  • Tune and validate time-series anomaly detection models for production monitoring.
  • Develop automation and operational tools using Python, Java, and React.
  • Troubleshoot Kubernetes, cloud, batch processing, and data pipeline issues.
  • Work with engineering and operations teams to improve system reliability and reduce manual work.
  • Support large-scale deployments and manage production risks such as configuration drift and blast radius.

Must-Have Skills

  • 8+ years of experience in SRE, DevOps, Platform Engineering, or Production Engineering.
  • Hands-on experience as an Incident Commander for P1/P2 incidents.
  • Strong experience with time-series anomaly detection models in production observability mandatory.
  • Strong production-level programming skills in:
    • Python
    • Java
    • React
  • Strong experience with SLI, SLO, and error budgets.
  • Hands-on observability experience with:
    • Prometheus
    • Grafana
    • OpenTelemetry
    • At least 2 log platforms such as Loki, Splunk, or Elasticsearch
  • Strong GCP experience.
  • Strong Kubernetes operational experience.
  • Experience with Rancher K3s.
  • Experience troubleshooting Apache Airflow and Tidal workflows/batch jobs.
  • Experience with production-scale distributed systems and on-call support.
Preferred Skills

  • Production Readiness Reviews / service launch experience.
  • BigQuery and PostgreSQL.
  • Chaos Engineering / fault injection.
  • TIC/Technical Incident Commander certification.
  • Healthcare, pharmacy, retail, or other high-availability environments.
  • LLM/GenAI for incident management, alert summarization, or runbook recommendations.
  • Kafka.
  • Istio / Envoy.
  • Terraform / Ansible.

Important: The Incident Commander experience and production time-series anomaly detection should be treated as hard requirements. A candidate who only has general monitoring/observability experience but has never worked with anomaly detection models would not be a strong fit.
Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search