Back to search
Pacer Group Linkedin · Posted 5d ago

SRE (Site Reliability Engineering) Lead

Woonsocket

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Job Title: SRE (Site Reliability Engineering) Lead

Location: Woonsocket, RI

Work Arrangement: Hybrid

Employment Type: Contract

Duration: 12 months

Domain: Healthcare / Retail Tech

Pay Rate: $41.00/hr W2 | $50.00/hr C2C

Application Deadline: September 15, 2026


SKILLS REQUIRED

Primary (Must-Have):

• 8+ years of Senior Software Engineering experience in SRE, DevOps, or Platform Engineering for distributed systems

• On-call Incident Commander (IC) experience leading P1/P2 incidents and tuning production time-series anomaly detection models

• Strong production programming proficiency in Python, React, and Java for operational tooling

• Hands-on design of SLIs, SLOs, and error budgets, plus deep observability experience (Prometheus, Grafana, OpenTelemetry, Loki/Splunk/Elasticsearch)

• Cloud platform expertise in GCP, Rancher K3s, advanced Kubernetes operations, and data pipeline observability (Airflow, Tidal)


Secondary (Good to Have):

• Experience owning Production Readiness Reviews, fault injection/chaos engineering, or holding TIC (Technical Incident Commander) certification

• Telemetry data analytics using SQL, Google BigQuery, and PostgreSQL

• Hands-on experience with streaming platforms (Kafka), service mesh (Istio, Envoy), and IaC (Terraform, Ansible)


POSITION OVERVIEW

Operating within the Platform Reliability and Infrastructure organization, this role leads Site Reliability Engineering initiatives for large-scale distributed production systems. Reporting to the Senior Manager of Infrastructure & Reliability, the SRE Lead will solve complex observability, incident response, and system availability challenges to ensure high availability, reduced blast radius, and continuous performance across business-critical customer and patient-facing applications.


ROLES & RESPONSIBILITIES

• Serve as an on-call Incident Commander (IC) for P1/P2 incidents, providing structured leadership updates and driving swift restoration of mission-critical services.

• Tune, validate, and manage time-series anomaly detection models within a production observability context to drive proactive incident mitigation.

• Design, implement, and maintain SLIs, SLOs, error budgets, and full-stack observability frameworks using Prometheus, Grafana, OpenTelemetry, and log aggregation suites.

• Write, deploy, and maintain production-quality operational tooling using Python, React, and Java across GCP, Rancher K3s, and Kubernetes environments.

• Diagnose and resolve complex workflow orchestration failures, batch processing issues, and scheduler performance bottlenecks within Airflow and Tidal data pipelines.


BENEFITS

Medical | Dental | Vision | 401(k)


EEOC Compliance:

We are an equal opportunity employer, and all qualified applicants will receive consideration for employment.


DISCLAIMER

AI Usage Policy: Pacer Group uses AI to assist in screening applications. Final hiring decisions are made by human recruiters based on qualifications and experience.

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search