Back to search
TestCrew | Quality Engineering & Software Testing Linkedin · Posted yesterday

Site Reliability Engineer

Riyadh

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

About the Role

TestCrew is seeking Site Reliability Engineers (SREs) to join an upcoming enterprise engagement in Al Ahsa. The successful candidates will be responsible for ensuring the availability, reliability, performance, and security of critical IT systems through proactive monitoring, automation, and disciplined incident response.

This is a hands-on operational role that requires close collaboration with development, operations, and security teams to maintain highly available services, streamline operations through automation, and drive continuous service improvement. Candidates must be willing to work onsite at the client location in Al Ahsa and participate in on-call or shift rotations as required.

Key Responsibilities
  • Monitor and manage enterprise infrastructure, applications, and services to ensure high availability, stability, and optimal performance.
  • Operate and maintain monitoring and observability platforms to detect incidents, analyze trends, and respond proactively to system issues.
  • Design and implement monitoring strategies, dashboards, alerts, and performance metrics to improve operational visibility.
  • Perform root cause analysis (RCA) for incidents and implement preventive measures to reduce recurring issues.
  • Automate operational tasks, deployments, and maintenance activities using scripting and infrastructure automation tools.
  • Manage CI/CD pipelines and support release management processes to enable reliable and efficient software deployments.
  • Optimize system performance while ensuring compliance with service level agreements (SLAs), operational standards, and best practices.
  • Support business continuity, disaster recovery, and operational resilience initiatives.
  • Implement security best practices, including system hardening, patch management, and secure operational procedures.
  • Collaborate with development, infrastructure, security, and IT operations teams to resolve complex technical issues.
  • Maintain operational documentation, runbooks, knowledge articles, and incident reports.
  • Participate in major incident management, post-incident reviews, and continuous improvement initiatives.
  • Provide on-call support and participate in shift rotations as required.
Required Qualifications
  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field.
  • 2–5 years of hands-on experience in Site Reliability Engineering (SRE), DevOps, Infrastructure Operations, or Production Support.
  • Experience managing enterprise monitoring and observability platforms such as:
  • Grafana
  • Prometheus
  • Datadog
  • Instana
  • Zabbix
  • or equivalent solutions
  • Strong troubleshooting and root cause analysis skills across infrastructure, applications, and networking.
  • Hands-on experience with scripting and automation using:
  • Python
  • Bash
  • PowerShell
  • Experience with automation and Infrastructure as Code (IaC) tools such as:
  • Ansible
  • Terraform
  • Practical knowledge of CI/CD tools such as:
  • Jenkins
  • GitLab CI
  • Azure DevOps
  • Good understanding of business continuity, disaster recovery, and operational risk management.
  • Knowledge of cybersecurity best practices, including system hardening, access control, vulnerability management, and patch management.
  • Familiarity with ITSM processes (Incident, Problem, and Change Management) and tools such as ServiceNow or Jira Service Management is an advantage.
  • Fluent in both Arabic and English, with strong written and verbal communication skills.
Preferred Qualifications
  • ITIL Foundation Certification.
  • Cloud platform experience (AWS, Azure, or Google Cloud Platform).
  • Experience with containerization and orchestration technologies such as Docker and Kubernetes.
  • Experience supporting large-scale enterprise or government environments.
Technical Skills
  • Site Reliability Engineering (SRE)
  • System Monitoring & Observability
  • Incident Management & Root Cause Analysis
  • Performance Monitoring & Optimization
  • DevOps & CI/CD
  • Infrastructure as Code (Terraform, Ansible)
  • Automation & Scripting (Python, Bash, PowerShell)
  • Linux & Windows Server Administration
  • Networking Fundamentals
  • Business Continuity & Disaster Recovery
  • Cybersecurity Operations
  • ITSM Processes
  • ServiceNow / Jira Service Management
  • Documentation & Operational Runbooks
Personal Attributes
  • Strong ownership mindset with a proactive, reliability-first approach.
  • Excellent analytical and problem-solving skills.
  • Ability to remain calm and methodical during critical incidents.
  • Strong communication and collaboration skills.
  • Ability to work effectively in a fast-paced enterprise environment.
  • Willingness to work onsite in Al Ahsa and participate in on-call or shift support as required.

  • Free. 20 seconds. No password. See every match in this search.

    Create a free Caio profile to unlock more results and save your role and location preferences.

    Unlock free search
    Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
    View Managed Job Search