Senior DevOps Engineer
Indexed description
We are seeking a Senior DevOps Engineer to design, build, and operate the cloud-native infrastructure powering our trading and data analytics platform. In this role, you will lead the implementation of automated CI/CD pipelines, container orchestration, infrastructure-as-code, and continuous operational monitoring across Linux-based cloud environments.
This position offers a clear trajectory to transition into steady-state DevOps and infrastructure ownership, playing a critical role in ensuring high system availability, zero-downtime deployments, and robust platform security.
- Design, build, and optimize automated CI/CD pipelines (e.g., GitHub Actions, GitLab CI, Jenkins) to ensure rapid, secure, and reliable software delivery across all microservices.
- Deploy, manage, and scale containerized applications using Docker and Kubernetes across development, staging, and production environments.
- Automate cloud infrastructure provisioning, configuration, and environment parity using modern IaC tools (e.g., Terraform, Pulumi).
- Maintain overall platform health and uptime. Actively participate in incident management, root-cause analysis, and systematic remediation to prevent recurring issues.
- Implement, maintain, and enhance enterprise logging, metric aggregation, and alerting stacks (e.g., Prometheus, Grafana, Loki, or cloud-native observability tooling).
- Enforce security best practices, Role-Based Access Controls (RBAC), network segmentation, and secrets management across all cloud environments.
3+ years of hands-on experience in DevOps, Site Reliability Engineering (SRE), or Infrastructure Engineering roles.
- Deep technical knowledge of Linux operating system internals, performance tuning, shell environments, and security hardening.
- Proven, production-level expertise managing Docker environments and Kubernetes clusters (CKA/CKAD certification is a plus).
- Strong proficiency in Python, Bash, or Go for writing maintenance scripts, custom tools, and pipeline automation.
- Production experience architecting and managing infrastructure on major public cloud providers (GCP preferred, AWS or Azure acceptable).
- olid grasp of core networking concepts, including TCP/IP, DNS, TLS/SSL termination, load balancers, VPC peering, and ingress controllers.
- Direct production experience with Terraform, Pulumi, or Ansible for declarative infrastructure management.
- Hands-on experience configuring and maintaining Prometheus, Grafana, GCP Cloud Monitoring, or Datadog.
- Background in supporting low-latency trading applications, high-frequency data pipelines, or real-time streaming architectures (e.g., Apache Kafka, WebSockets).
- Experience designing and operating multi-region, disaster-recovery ready, or multi-zone high-availability cloud deployments.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search