Site Reliability Engineer
Indexed description
Job Title: Site Reliability Engineer (SRE)
Work Setup: Hybrid (3 Days Onsite, 2 Days Remote)
Location: Eton Centris, Quezon City, Philippines
Employment Type: Full-Time
Job Summary
We are seeking a highly skilled Site Reliability Engineer (SRE) to join our growing technology team. The ideal candidate will be responsible for maintaining the reliability, availability, scalability, and performance of critical applications and infrastructure. This role combines software engineering and systems administration principles to build and operate resilient, automated, and highly available platforms.
The successful candidate will work closely with development, infrastructure, security, and operations teams to improve system reliability, optimize performance, and enhance operational efficiency through automation and continuous improvement initiatives.
Key Responsibilities
- Design, implement, and maintain highly available and scalable infrastructure and application environments.
- Monitor system health, application performance, and service availability using industry-standard monitoring and observability tools.
- Develop and maintain automation scripts, tools, and workflows to improve operational efficiency.
- Manage incident response, troubleshooting, root cause analysis (RCA), and post-incident reviews.
- Collaborate with development teams to improve application reliability, deployment processes, and operational readiness.
- Establish and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).
- Implement infrastructure-as-code (IaC) solutions for consistent and repeatable deployments.
- Support CI/CD pipelines and deployment automation initiatives.
- Perform capacity planning, performance tuning, and proactive system optimization.
- Ensure compliance with security, governance, and operational best practices.
- Create and maintain technical documentation, operational runbooks, and disaster recovery procedures.
- Participate in on-call support and incident management activities as required.
Required Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- Minimum of 3 years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, Systems Engineering, or Infrastructure Operations.
- Experience supporting production environments with high availability and uptime requirements.
- Strong understanding of Linux and/or Windows server administration.
- Experience with cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform (GCP).
- Experience with containerization technologies such as Docker and orchestration platforms such as Kubernetes.
- Hands-on experience with CI/CD tools such as Azure DevOps, GitHub Actions, Jenkins, GitLab CI/CD, or similar.
- Experience with infrastructure-as-code tools such as Terraform, CloudFormation, or Ansible.
- Familiarity with monitoring and observability platforms such as Prometheus, Grafana, Datadog, New Relic, Dynatrace, Splunk, or ELK Stack.
- Strong scripting and automation skills using Python, Bash, PowerShell, or similar languages.
- Experience in troubleshooting complex production incidents and conducting root cause analysis.
- Knowledge of networking concepts, DNS, load balancing, firewalls, and security best practices.
Preferred Qualifications
- Experience working in enterprise or cloud-native environments.
- Familiarity with microservices architecture and distributed systems.
- Experience implementing reliability engineering practices, including SLOs, SLIs, and error budgets.
- Cloud certifications (AWS, Azure, or GCP) are an advantage.
- Experience with security and compliance frameworks.
- Knowledge of disaster recovery and business continuity planning.
Technical Skills
- Linux / Windows Administration
- AWS, Azure, or GCP
- Kubernetes & Docker
- Terraform, Ansible, CloudFormation
- CI/CD Pipelines
- Git Version Control
- Monitoring & Observability Tools
- Python, Bash, PowerShell
- Networking & Security Fundamentals
- Incident Management & Root Cause Analysis
- Infrastructure Automation
- High Availability & Disaster Recovery
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search