Senior Site Reliability Engineer
Indexed description
You will be responsible for designing, building, automating, and maintaining secure cloud infrastructure that enables engineering teams to deliver software quickly and reliably.
You will work closely with software engineers to implement Infrastructure as Code, improve CI/CD pipelines, strengthen security, enhance observability, and ensure high platform availability.
Key Responsibilities
- Design, build, and manage scalable AWS cloud infrastructure.
- Develop and maintain Infrastructure as Code using Terraform.
- Build, optimize, and maintain Kubernetes clusters (AWS EKS).
- Implement and improve CI/CD pipelines for automated application deployments.
- Manage containerized workloads using Docker and Kubernetes.
- Implement GitOps workflows using ArgoCD.
- Design monitoring, logging, and alerting solutions using Prometheus, Grafana, OpenTelemetry, and related tools.
- Improve system reliability, availability, and disaster recovery processes.
- Implement cloud security best practices across infrastructure and deployment pipelines.
- Automate infrastructure provisioning and operational workflows.
- Perform infrastructure troubleshooting, root cause analysis, and incident response.
- Collaborate with engineering teams to improve deployment velocity and developer experience.
- Optimize cloud costs while maintaining platform performance.
- Maintain documentation for infrastructure, deployment processes, and operational procedures.
Required Qualification
- 5+ years of experience in DevOps, Site Reliability Engineering, or Platform Engineering.
- Strong hands-on experience with AWS, Kubernetes (EKS), Docker, Terraform, and Python.
- Experience building CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or similar tools.
- Experience developing automation scripts and infrastructure tooling using Python.
- Solid understanding of Linux, networking, cloud security, and Infrastructure as Code.
- Experience with monitoring and observability tools such as Prometheus, Grafana, and OpenTelemetry.
- Strong scripting skills using Python and/or Bash.
- Excellent problem-solving and communication skills.
Nice To Have
- Experience with Helm and ArgoCD.
- Knowledge of OpenTelemetry, distributed tracing, and GitOps.
- Experience with security tools such as Trivy or Falco.
- AWS Certifications.
- Experience working in fintech, AI, or other highly regulated environments.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search