IT Operations Engineer (DevOps & Production Services)
Indexed description
The successful candidate will drive operational excellence through automation, observability, incident prevention, and continuous improvement initiatives, helping maintain highly available and resilient production environments.
Key Responsibilities
- Ensure the stability, availability, performance, and security of production applications and infrastructure.
- Collaborate with Development, Infrastructure, Security, and Support teams to integrate operational requirements into solution design and delivery.
- Monitor and proactively improve system health, performance, and reliability through observability and automation practices.
- Support incident, problem, and change management processes, driving root cause analysis and preventive actions.
- Contribute to infrastructure automation and Infrastructure as Code (IaC) initiatives.
- Participate in deployment, release, and environment management activities within CI/CD pipelines.
- Promote operational best practices aligned with Agile, DevOps, SRE, and IT Service Management principles.
- Continuously identify opportunities to improve operational efficiency, resilience, and service quality.
- 2+ years of experience in IT Operations, DevOps, Site Reliability Engineering (SRE), Infrastructure, or related disciplines.
- Experience working in Agile and DevOps environments.
- Proven experience supporting business-critical production applications and infrastructure.
- Solid understanding of:
- IT Infrastructure Management
- IT Service Management (ITSM)
- Incident, Problem, and Change Management
- IT Risk Management
- Cybersecurity fundamentals
- Experience implementing operational monitoring, observability, and automation solutions.
- Strong analytical and problem-solving skills.
- Ability to work effectively in cross-functional and globally distributed teams.
- Docker
- Kubernetes
- OpenShift
- Helm
- GitLab
- Jenkins
- Artifactory
- SonarQube
- Ansible
- Mattermost
- Elastic Stack (ELK)
- Grafana
- Dynatrace
- Linux
- Windows Server
- English: Fluent/Professional Proficiency (Mandatory)
- Experience with cloud platforms (AWS, Azure, or Google Cloud Platform).
- Knowledge of CI/CD pipeline design and management.
- Experience with scripting and automation using Shell, Python, or PowerShell.
- Familiarity with security and compliance requirements in enterprise environments.
- Exposure to high-availability and disaster recovery architectures.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search