Back to search
SRE Linkedin · Posted yesterday

Site Reliability Engineer

Puerto Rico

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Job Title: Site Reliability Engineer (SRE)

Location: Puerto Rico (Remote)

Position Overview

We are seeking a highly skilled Site Reliability Engineer (SRE) to ensure the reliability, scalability, performance, and operational excellence of mission-critical production environments. The ideal candidate will combine strong software engineering principles with infrastructure expertise to build resilient platforms, automate operations, enhance observability, and minimize operational toil.

As an SRE, you will collaborate closely with Platform Engineering, DevOps, Infrastructure, Security, and Development teams to deliver highly available, secure, and scalable services across hybrid cloud and on-premises environments.


Key Responsibilities

Reliability & System Performance

  • Ensure high availability, scalability, and performance of production environments.
  • Define, implement, and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
  • Perform proactive capacity planning, performance tuning, and system optimization.
  • Lead Root Cause Analysis (RCA) for production incidents and implement preventive solutions.

Automation & Platform Engineering

  • Design and develop automation to eliminate manual operational tasks.
  • Build reusable Infrastructure as Code (IaC) using Terraform, Ansible, Pulumi, or similar tools.
  • Automate deployments, provisioning, patching, configuration management, and operational workflows.
  • Improve developer productivity through self-service platform capabilities.

CI/CD & Release Engineering

  • Design, build, and optimize CI/CD pipelines.
  • Support automated deployments using:
  • GitHub Actions
  • Jenkins
  • GitLab CI
  • Azure DevOps
  • Ensure safe, repeatable, and low-risk software releases.
  • Implement deployment strategies such as Blue/Green and Canary deployments.

Monitoring & Observability

  • Design enterprise-grade monitoring and alerting solutions.
  • Build dashboards and telemetry using:
  • Prometheus
  • Grafana
  • ELK Stack
  • Datadog
  • CloudWatch
  • Splunk (preferred)
  • Improve application and infrastructure observability using logs, metrics, and distributed tracing.

Incident Response & Reliability Engineering

  • Participate in 24x7 on-call rotations.
  • Lead production incident response and service restoration.
  • Drive post-incident reviews and continuous improvement initiatives.
  • Reduce Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).

Cloud & Infrastructure Management

  • Manage cloud infrastructure across:
  • AWS
  • Microsoft Azure
  • Google Cloud Platform (GCP)
  • Support hybrid environments including:
  • VMware
  • Nutanix
  • Deploy and manage highly available infrastructure services.

Container & Kubernetes Platform

  • Deploy and manage Kubernetes clusters.
  • Support:
  • Kubernetes
  • Amazon EKS
  • Azure AKS
  • Amazon ECS
  • Troubleshoot containerized workloads and optimize cluster performance.


Share resume to: [email protected]

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search