Site Reliability Engineer
Indexed description
TestCrew is seeking Site Reliability Engineers (SREs) to join an upcoming enterprise engagement in Al Ahsa. The successful candidates will be responsible for ensuring the availability, reliability, performance, and security of critical IT systems through proactive monitoring, automation, and disciplined incident response.
This is a hands-on operational role that requires close collaboration with development, operations, and security teams to maintain highly available services, streamline operations through automation, and drive continuous service improvement. Candidates must be willing to work onsite at the client location in Al Ahsa and participate in on-call or shift rotations as required.
Key Responsibilities- Monitor and manage enterprise infrastructure, applications, and services to ensure high availability, stability, and optimal performance.
- Operate and maintain monitoring and observability platforms to detect incidents, analyze trends, and respond proactively to system issues.
- Design and implement monitoring strategies, dashboards, alerts, and performance metrics to improve operational visibility.
- Perform root cause analysis (RCA) for incidents and implement preventive measures to reduce recurring issues.
- Automate operational tasks, deployments, and maintenance activities using scripting and infrastructure automation tools.
- Manage CI/CD pipelines and support release management processes to enable reliable and efficient software deployments.
- Optimize system performance while ensuring compliance with service level agreements (SLAs), operational standards, and best practices.
- Support business continuity, disaster recovery, and operational resilience initiatives.
- Implement security best practices, including system hardening, patch management, and secure operational procedures.
- Collaborate with development, infrastructure, security, and IT operations teams to resolve complex technical issues.
- Maintain operational documentation, runbooks, knowledge articles, and incident reports.
- Participate in major incident management, post-incident reviews, and continuous improvement initiatives.
- Provide on-call support and participate in shift rotations as required.
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field.
- 2–5 years of hands-on experience in Site Reliability Engineering (SRE), DevOps, Infrastructure Operations, or Production Support.
- Experience managing enterprise monitoring and observability platforms such as:
- Grafana
- Prometheus
- Datadog
- Instana
- Zabbix
- or equivalent solutions
- Strong troubleshooting and root cause analysis skills across infrastructure, applications, and networking.
- Hands-on experience with scripting and automation using:
- Python
- Bash
- PowerShell
- Experience with automation and Infrastructure as Code (IaC) tools such as:
- Ansible
- Terraform
- Practical knowledge of CI/CD tools such as:
- Jenkins
- GitLab CI
- Azure DevOps
- Good understanding of business continuity, disaster recovery, and operational risk management.
- Knowledge of cybersecurity best practices, including system hardening, access control, vulnerability management, and patch management.
- Familiarity with ITSM processes (Incident, Problem, and Change Management) and tools such as ServiceNow or Jira Service Management is an advantage.
- Fluent in both Arabic and English, with strong written and verbal communication skills.
- ITIL Foundation Certification.
- Cloud platform experience (AWS, Azure, or Google Cloud Platform).
- Experience with containerization and orchestration technologies such as Docker and Kubernetes.
- Experience supporting large-scale enterprise or government environments.
- Site Reliability Engineering (SRE)
- System Monitoring & Observability
- Incident Management & Root Cause Analysis
- Performance Monitoring & Optimization
- DevOps & CI/CD
- Infrastructure as Code (Terraform, Ansible)
- Automation & Scripting (Python, Bash, PowerShell)
- Linux & Windows Server Administration
- Networking Fundamentals
- Business Continuity & Disaster Recovery
- Cybersecurity Operations
- ITSM Processes
- ServiceNow / Jira Service Management
- Documentation & Operational Runbooks
- Strong ownership mindset with a proactive, reliability-first approach.
- Excellent analytical and problem-solving skills.
- Ability to remain calm and methodical during critical incidents.
- Strong communication and collaboration skills.
- Ability to work effectively in a fast-paced enterprise environment.
- Willingness to work onsite in Al Ahsa and participate in on-call or shift support as required.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search