DevOps & Site Reliability Engineer
Indexed description
Role Overview
Aivar Innovations is looking for a skilled DevOps & Site Reliability Engineer to join our Platform Engineering team. In this role you will own the reliability, scalability, and operational excellence of our cloud-native infrastructure. You will partner closely with product engineering teams to design and operate CI/CD pipelines, Kubernetes-based workloads on AWS, and observability stacks — ensuring that Aivar's accelerator platforms run with the highest possible uptime and performance.
This is a hands-on role with significant ownership. You will be expected to drive automation-first thinking, reduce toil, and champion SRE principles across the organisation.
Requirements
KEY RESPONSIBILITIES
Infrastructure & Platform Operations
- Design, provision, and maintain production-grade AWS infrastructure using Infrastructure-as-Code (Terraform / CloudFormation).
- Manage and optimise multi-cluster Kubernetes environments (EKS) — including cluster upgrades, node scaling, network policies, and security hardening.
- Administer and tune AWS-managed services: RDS, ElastiCache, MSK (Kafka), S3, CloudFront, Route 53, and VPC networking.
- Implement and own cost-optimisation strategies — right-sizing, spot usage, savings plans, and tagging governance.
- Build, maintain, and improve CI/CD pipelines using GitHub Actions, Jenkins, or equivalent tooling.
- Drive automation of operational tasks — provisioning, patching, scaling events, and incident runbooks — to minimise manual intervention.
- Champion GitOps workflows using ArgoCD or Flux for Kubernetes application delivery.
- Develop and maintain Helm charts and Kustomize overlays for consistent, repeatable deployments across environments.
- Define and track SLOs, SLIs, and error budgets for production services.
- Build and operate observability platforms using Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent tooling.
- Lead incident response efforts — on-call participation, post-incident reviews, and long-term remediation tracking.
- Perform chaos engineering experiments and game days to proactively surface reliability risks.
- Enforce security best practices across the platform — IAM least-privilege, secrets management (AWS Secrets Manager / Vault), container image scanning.
- Collaborate with the security team on vulnerability management and patching cadence.
- 3 – 5 years of hands-on experience in a DevOps, Platform Engineering, or Site Reliability Engineering role.
- Proven track record managing production Kubernetes clusters at scale — EKS strongly preferred.
- Deep working knowledge of AWS services, including compute (EC2, ECS, Lambda), networking (VPC, ALB, NLB, Transit Gateway), storage, and managed databases.
- Strong proficiency in at least one IaC tool: Terraform (preferred), AWS CDK, or CloudFormation.
- Solid scripting ability in Python, Bash, or Go for automation and tooling development.
- AWS Certified DevOps Engineer – Professional or AWS Certified Solutions Architect certification.
- Certified Kubernetes Administrator (CKA) or Certified Kubernetes Security Specialist (CKS).
- Experience with service mesh technologies such as Istio or Linkerd.
- Prior experience operating AI/ML or data-intensive workloads on Kubernetes (GPU node pools, Karpenter, etc.).
- Exposure to multi-cloud or hybrid-cloud environments.
- Learn from Experts: Work directly with former AWS leaders and AI pioneers.
- Direct Ownership: Lead high-impact "greenfield" projects from concept to global launch.
- Modern Tech: Master the latest Generative AI frameworks and cloud-native architectures.
- Real-World Impact: Build mission-critical systems used by major global enterprises.
- Rapid Growth: Scale your career quickly in a high-speep
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search