DevOps Engineer
Indexed description
Working with software engineers, security, and platform teams, the role improves deployment velocity and operational resilience across the stack. Success means faster, safer releases, clear ownership of production systems, and measurable improvements in availability, latency, and incident response.
Key Responsibilities
- Design and maintain AWS infrastructure using Terraform, including VPCs, IAM, ECS or EKS, RDS, S3, and CloudWatch
- Build and improve CI/CD pipelines with GitHub Actions, GitLab CI, or Jenkins for automated testing, deployment, rollback, and environment promotion
- Operate Kubernetes workloads in production, including cluster configuration, Helm releases, autoscaling, networking, and resource management
- Implement observability with Prometheus, Grafana, ELK or OpenSearch, and distributed tracing to identify performance and reliability issues
- Automate operational workflows with Python, Go, or Bash, reducing repetitive work across provisioning, releases, incident response, and access management
- Define reliability practices including health checks, alerting, runbooks, backup validation, disaster recovery procedures, and capacity planning
- Participate in on-call rotations, lead technical incident response, and drive post-incident remediation through measurable engineering improvements
- 3–8 years of experience in DevOps, site reliability engineering, cloud infrastructure, or platform engineering
- Hands-on experience operating production workloads in AWS and managing infrastructure as code with Terraform or an equivalent tool
- Strong Kubernetes experience, including deployments, services, ingress, Helm, autoscaling, and troubleshooting containerized applications
- Proficiency with Linux administration, networking fundamentals, Git, and at least one scripting or programming language such as Python, Go, or Bash
- Experience building CI/CD pipelines and applying automated testing, secrets management, security controls, and safe deployment strategies
- Working knowledge of observability, incident management, high availability, disaster recovery, and common reliability metrics such as SLOs, SLIs, and error budgets
- Bachelor’s degree in computer science, information technology, engineering, or a related field; equivalent practical experience is accepted. Bonus: experience with Argo CD, service meshes, AWS certifications, compliance frameworks, or multi-region infrastructure
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search