Cloud DevOps Engineer - Azure/AWS
Indexed description
- Job Details
Department / Business Line: Technology Services
Reports To: Infrastructure Manager
Employment Type: Permanent
Work Location: Maadi, Degla - On-Site
Grade: As per organizational structure
- Job Purpose
The role focuses on infrastructure automation, cloud operations, continuous integration and deployment, system monitoring, security, availability, disaster recovery, and cost optimization across development, UAT, and production environments.
- Key Accountabilities & Deliverables
- Provisioned, configured, and fully documented Development, UAT, and Production cloud environments
- Infrastructure-as-Code (IaC) templates using Terraform, ARM, or Bicep
- Approved and version-controlled cloud architecture and infrastructure diagrams
- Regular cloud cost monitoring and optimization reports
- Fully automated CI/CD pipelines covering build, testing, deployment, and release processes
- Automated rollback and disaster recovery procedures
- System and application monitoring dashboards
- SLA/SLO-based alerting and notification configurations
- Monthly system uptime and availability reports, including agreed availability targets such as 99.9%
- Incident Root Cause Analysis (RCA) reports and corrective action plans
- Cloud security baselines aligned with applicable security and governance standards, including ISO 27001 and relevant national cybersecurity requirements
- Vulnerability assessment and patch compliance reports
- IAM access control matrices, including roles, permissions, and access reviews
- Audit logs and compliance evidence repositories
- Cloud backup policies, configurations, monitoring, and recovery procedures
- Technical documentation for infrastructure, deployments, operational procedures, and troubleshooting
- Key Responsibilities
- CI/CD Pipeline Management
- Design, implement, maintain, and continuously improve CI/CD pipelines using Azure DevOps, Jenkins, GitLab CI, or equivalent tools
- Automate build, testing, deployment, and release workflows to minimize manual intervention and deployment risks
- Implement deployment strategies such as Blue/Green, Rolling, and Canary deployments, where applicable
- Monitor pipeline performance and troubleshoot build, deployment, and release failures
- Implement automated quality gates, security checks, and approval processes within deployment pipelines
- Support development teams in adopting DevOps best practices
- Design, implement, and maintain cloud infrastructure using Terraform, Ansible, ARM Templates, or Bicep
- Develop reusable and standardized infrastructure modules and automation scripts
- Maintain infrastructure code in version-controlled repositories
- Ensure infrastructure changes follow appropriate change management, review, and approval processes
- Promote Infrastructure-as-Code practices across the Technology Services team
- Automate repetitive operational activities to improve efficiency and reliability
- Provision, configure, manage, and maintain cloud resources across Microsoft Azure and/or AWS
- Manage cloud compute, storage, networking, identity, security, and supporting services
- Maintain separate and properly governed Dev, UAT, and Production environments
- Monitor cloud resource utilization, performance, capacity, and availability
- Identify opportunities for cloud cost optimization and resource efficiency
- Ensure cloud environments comply with organizational security, governance, and operational standards
- Build, manage, and maintain containerized applications using Docker
- Deploy and operate workloads on Kubernetes environments
- Manage Kubernetes configurations, deployments, services, ingress, secrets, and resource utilization
- Support container security, scalability, availability, and performance optimization
- Troubleshoot container and Kubernetes-related operational issues
- Implement and maintain infrastructure and application monitoring solutions
- Develop dashboards and operational metrics to provide visibility into system health and performance
- Configure proactive alerts based on SLA, SLO, performance, and availability requirements
- Monitor production environments and respond to operational incidents
- Participate in incident management, problem management, and service restoration activities
- Conduct Root Cause Analysis (RCA) for significant incidents and implement preventive actions
- Support continuous improvement of system reliability and availability
- Implement and maintain cloud security baselines and security best practices
- Manage Identity and Access Management (IAM), roles, permissions, and access controls
- Support vulnerability management, security patching, and remediation activities
- Maintain audit logs and technical evidence required for internal and external audits
- Ensure cloud infrastructure follows applicable organizational security policies and regulatory requirements
- Support security assessments and compliance initiatives aligned with standards such as ISO 27001 and relevant national cybersecurity requirements
- Design and maintain cloud backup and recovery solutions
- Ensure backup configurations meet business and operational requirements
- Regularly monitor backup jobs and investigate failures
- Develop and maintain disaster recovery procedures and automation
- Participate in backup restoration and disaster recovery testing
- Maintain documentation for recovery procedures, dependencies, and recovery objectives
- Maintain accurate and up-to-date infrastructure and architecture documentation
- Document operational procedures, deployment processes, recovery procedures, and troubleshooting guides
- Maintain version-controlled technical documentation and configuration repositories
- Share knowledge and provide technical guidance to infrastructure, development, and support teams
- Required Technical Skills
- Azure DevOps
- Jenkins
- GitLab CI/CD
- Git and version control
- Automated build, testing, deployment, and release management
- Terraform
- Ansible
- ARM Templates
- Bicep
- Infrastructure automation and configuration management
- Microsoft Azure and/or AWS
- Cloud compute, storage, networking, IAM, monitoring, and security services
- Cloud cost management and optimization
- Docker
- Kubernetes
- Container deployment and management
- Python
- Bash
- PowerShell
- Cloud monitoring and logging platforms
- Dashboard development
- Alerting and incident management
- Performance and availability monitoring
- Cloud IAM
- Vulnerability and patch management
- Security baselines
- Audit logging
- ISO 27001 awareness
- Relevant national cybersecurity and compliance requirements
- Qualifications & Experience
- Bachelor's degree in Computer Science, Information Technology, Computer Engineering, or a related Engineering/Technology discipline
- Microsoft Certified: DevOps Engineer Expert
- AWS Certified DevOps Engineer - Professional
- Certified Kubernetes Administrator (CKA) - preferred
- Relevant Terraform, cloud, or security certifications are an advantage
- 3-5 years of proven experience in DevOps, Cloud Engineering, Platform Engineering, Infrastructure Engineering, or a closely related role
- Proven experience designing and implementing CI/CD pipelines
- Proven hands-on experience with cloud infrastructure and Infrastructure-as-Code
- Experience supporting production environments and resolving infrastructure and deployment incidents
- Practical experience with containerization and Kubernetes is highly desirable
- Core Competencies
- Strong problem-solving and analytical skills
- Excellent troubleshooting and incident-resolution capabilities
- Strong understanding of DevOps and automation principles
- Ability to work effectively across development, infrastructure, security, and business teams
- Strong focus on automation, reliability, security, and operational excellence
- Ability to work under pressure and manage production incidents effectively
- Strong documentation and communication skills
- Continuous learning mindset and ability to adapt to emerging cloud and DevOps technologies
- Language Requirements
- Arabic: Professional proficiency
- English: Professional proficiency
- Job Dimensions
Financial Responsibility: None / No direct financial responsibility
- Key Performance Indicators (KPIs)
- CI/CD pipeline reliability and success rate
- Deployment frequency and deployment lead time
- Reduction in manual deployment activities
- Infrastructure provisioning and automation efficiency
- Cloud availability and uptime targets
- Incident response and resolution times
- Number and quality of completed RCA reports
- Backup success and recovery test results
- Vulnerability and patch compliance
- Cloud cost optimization and resource utilization
- Compliance with security and governance requirements
- Accuracy and completeness of technical documentation
- Job Requirements
- Ability to work on-site from the Maadi, Degla office
- Willingness to participate in production support and incident management activities when required
- Ability to work collaboratively with cross-functional technology teams
- Strong commitment to security, reliability, automation, and continuous improvement
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search