DevOps Specialist
Indexed description
Senior DevOps Engineer
About the Role
We are looking for a highly experienced Senior DevOps Engineer with a strong technical background and a proven track record of designing, implementing, and operating complex, scalable technology environments.
The ideal candidate will be a key technical reference within the engineering organisation, contributing to architecture decisions, establishing technical standards, driving automation, and ensuring the reliability, scalability, security, and efficiency of our platforms.
Mandatory Requirements
- 10+ years of professional experience in DevOps, Platform Engineering, SRE, Cloud Infrastructure, or a closely related field.
- Proven experience working with complex, distributed, and highly scalable environments.
- Strong experience designing and implementing cloud architectures, preferably AWS, Azure, or GCP.
- Advanced knowledge of Kubernetes and containerised environments.
- Extensive experience with Infrastructure as Code (IaC), particularly Terraform.
- Strong expertise in CI/CD, including the design and optimisation of deployment pipelines.
- Advanced Linux and systems administration knowledge.
- Strong understanding of networking, security, IAM, authentication, and access management.
- Hands-on experience with observability, including monitoring, logging, alerting, tracing, and incident management.
- Strong experience with Git and modern software development workflows.
- Solid scripting/programming skills in languages such as Python, Bash, Go, or similar.
- Experience designing systems for high availability, scalability, resilience, and disaster recovery.
- Strong understanding of DevSecOps principles and the integration of security into the development and deployment lifecycle.
- Experience implementing and maintaining production-grade platforms and environments.
Key Responsibilities
- Define and evolve the technical direction of the DevOps and platform environment.
- Design scalable, secure, highly available, and resilient infrastructure architectures.
- Establish and maintain engineering standards, best practices, and technical guidelines.
- Drive the adoption of automation across infrastructure, deployments, testing, and operational processes.
- Design, implement, and continuously improve CI/CD pipelines.
- Own and improve the organization's cloud and infrastructure strategy.
- Lead technical initiatives involving Kubernetes, cloud platforms, IaC, observability, security, and automation.
- Identify technical risks and proactively implement solutions to improve system reliability and performance.
- Drive improvements in system availability, scalability, security, and operational efficiency.
- Participate in architectural decisions and technical evaluations of new technologies.
- Collaborate closely with Software Engineering, Architecture, Security, and Product teams.
- Act as a technical point of reference for complex DevOps and infrastructure challenges.
- Mentor and support other engineers, promoting knowledge sharing and strong engineering practices.
- Participate in incident analysis, root-cause investigations, and the implementation of long-term corrective actions.
Technical Expertise
The successful candidate is expected to have strong expertise across several of the following areas:
Cloud & Infrastructure
- AWS, Azure, or GCP
- VPC/VNet, networking, load balancing, DNS, IAM
- High availability and disaster recovery
- Cloud security and cost optimization
Containers & Orchestration
- Kubernetes
- Docker
- Helm
- Container security and orchestration
Infrastructure as Code & Automation
- Terraform
- Ansible or similar configuration management tools
- Infrastructure automation and provisioning
- GitOps practices
CI/CD
- Jenkins, GitHub Actions, GitLab CI, Azure DevOps, or similar
- Automated build, test, deployment, and release processes
- Deployment strategies such as blue/green, canary, and rolling deployments
Observability & Reliability
- Prometheus, Grafana, ELK/OpenSearch, Datadog, New Relic, or similar
- Metrics, logs, traces, alerting, and dashboards
- SLI/SLO concepts
- Incident response and root-cause analysis
Security
- IAM and secrets management
- Network security
- Vulnerability management
- Secure CI/CD pipelines
- DevSecOps practices
Development & Automation
- Strong Bash and/or Python
- Working knowledge of Go or another programming language is a plus
- Git and Git-based development workflows
Leadership & Collaboration
- Ability to provide technical direction and influence engineering decisions.
- Comfortable owning complex initiatives from definition through implementation.
- Ability to establish technical standards and drive their adoption across teams.
- Strong problem-solving and decision-making skills.
- Ability to communicate complex technical concepts clearly to both technical and non-technical stakeholders.
- Strong mentoring and knowledge-sharing capabilities.
- Comfortable challenging existing approaches and proposing pragmatic improvements.
- Ability to balance technical excellence, business priorities, scalability, and delivery timelines.
Nice to Have
- Experience in SRE or Platform Engineering environments.
- Experience operating large-scale production platforms.
- Experience with GitOps tools such as Argo CD or Flux.
- Experience with service mesh technologies.
- Experience with Kafka or other distributed messaging systems.
- Experience with database infrastructure and performance optimization.
- Relevant cloud or Kubernetes certifications.
- Experience working in international, distributed engineering teams.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search