Senior Site Reliability Engineer (Azure)
Indexed description
Location: Abu Dhabi, UAE
Employment: 12-months initial (contract to perm)
We are seeking a hands-on Senior DevOps / Site Reliability Engineer to build and operate the delivery and runtime foundations for a portfolio of modern enterprise applications, workflow platforms and AI-enabled products.
This is not primarily an infrastructure-administration role. You will work directly with software engineers to create reliable, secure and automated paths from code to production.
You will own deployment automation, runtime reliability, observability, infrastructure-as-code and operational readiness across applications that integrate with critical enterprise systems.
What You Will Own
Platform & Infrastructure
* Build and maintain cloud infrastructure using Infrastructure as Code.
* Design secure, repeatable environments across development, test, staging and production.
* Manage containerized workloads and Kubernetes-based deployments where appropriate.
* Define standard application deployment patterns for backend, frontend and AI services.
* Implement secure secrets and configuration management.
* Support network, identity and connectivity requirements for enterprise integrations.
CI/CD & Developer Productivity
* Build automated CI/CD pipelines.
* Standardize build, test, security scanning and deployment processes.
* Automate environment provisioning and configuration.
* Reduce manual deployment steps and production configuration drift.
* Work closely with engineering teams to improve release frequency and reliability.
Reliability & Observability
* Establish logging, metrics, tracing and alerting.
* Define service-level indicators and operational thresholds.
* Build dashboards for system health and application performance.
* Implement incident-response and production-support practices.
* Design for graceful degradation, retries, failover and recovery.
* Lead root-cause analysis of production incidents.
Security & Operational Controls
* Implement least-privilege access and secure deployment patterns.
* Support auditability of infrastructure and production changes.
* Integrate security checks into delivery pipelines.
* Work with security and infrastructure teams to meet enterprise control requirements.
Resilience
* Support business continuity and disaster-recovery design.
* Define backup, restore and recovery procedures.
* Test operational recovery rather than relying solely on documented plans.
Required Experience
* 6+ years in DevOps, SRE, platform engineering or cloud infrastructure.
* Strong production experience with Azure, AWS or GCP; Azure strongly preferred.
* Docker and Kubernetes.
* Infrastructure as Code using Terraform, Bicep, Pulumi or equivalent.
* CI/CD using Azure DevOps, GitHub Actions, GitLab CI or similar.
* Strong Linux and networking fundamentals.
* Observability tooling and distributed-system troubleshooting.
* Secure secrets, identity and access-management patterns.
* Production incident-management experience.
* Scripting/programming capability in Python, Go, Bash or equivalent.
Strong Advantage
* Azure Kubernetes Service.
* Azure Service Bus, API Management, Key Vault and related Azure services.
* Enterprise integration platforms.
* SAP-connected environments.
* AI/LLM application deployment.
* Regulated or government environments.
* High-availability and disaster-recovery architecture.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search