Senior DevOps (GCP)
Indexed description
We are hiring two senior engineers who are equally at home in DevOps, cloud architecture and SRE to own this platform end to end — deploy paths, infrastructure as code, Kubernetes, secrets and identity, reliability and disaster recovery — and to set the standard the rest of engineering follows.
Key Responsibilities:
- Deploy paths: Converge our Cloud Run and GKE deployments onto one reviewed pattern — GitHub Actions with keyless Workload Identity Federation, additive configuration semantics, every change through a pull request — working repository by repository with the teams that own them.
- Infrastructure as code: Grow our Terraform estate (Google, Google-beta, MongoDB Atlas and Alibaba providers) from foundations — VPC, IAM, security, zero-trust access — through workloads to Kubernetes, with remote state, plan-in-PR, apply-on-merge and a drift policy people actually follow.
- Kubernetes: GKE operations end to end: node pools, Workload Identity, least-privilege RBAC, Helm charts, Istio / Gateway API ingress, backup and restore, zero-downtime upgrades.
- Secrets and identity: Drive the move from inline configuration to Secret Manager references; own service-account and IAM hygiene; make the audit trail the answer to "who changed what, when, approved by whom".
- Reliability: Define SLOs for the services that matter, build alerting and dashboards (Cloud Monitoring, OpenTelemetry), run incident response, write blameless postmortems and turn each one into an engineering change.
- Disaster recovery: Keep the Alibaba Cloud DR footprint — ACK, VPC, KMS, OSS — provisioned by Terraform and proven by regular exercise.
- The AI DevOps agent: Our agent proposes IAM grants, restarts, configuration and Cloud Run changes; humans approve; the platform executes and records. You will shape what it may propose, review the pull requests it opens and harden its guardrails. You are the human in that loop.
- 6+ years building and running production infrastructure, 3+ on Google Cloud at depth: GKE, Cloud Run, IAM, Secret Manager, VPC and networking, Artifact Registry, Cloud Build, Cloud Logging and Monitoring.
- Terraform in anger module design, state management, imports, refactors without destroys, provider upgrades.
- Kubernetes in production Helm, RBAC, network policy, an ingress or mesh layer (Istio or Gateway API), and troubleshooting from kubectl down to the node.
- CI/CD you have designed, not just used: GitHub Actions (reusable workflows, OIDC / WIF, environments with required reviewers) and Cloud Build.
- SRE practice: SLOs and error budgets, on-call, incident command, postmortems, capacity and cost.
- Security as a habit: least privilege, secrets never in code or logs, audit trails — and the judgement to say "not without approval".
- Linux, networking and scripting: Bash and Python are daily tools.
- Written clarity: Our change process runs on pull requests and written decisions. You can state what you measured, what you changed and what you could not verify.
- Comfortable working remotely across time zones, with async communication as the default.
- HashiCorp Terraform Associate (or higher) certification.
- Google Cloud Professional Cloud DevOps Engineer or Professional Cloud Architect certification.
- Alibaba Cloud (ACK, OSS, KMS, VPC) — our DR platform.
- MongoDB Atlas, PostgreSQL, Redis or Kafka operations.
- Zero-trust network access and OIDC identity providers.
- Regulated-environment experience: data residency, financial-services controls, audit readiness.
- Experience running or reviewing LLM-based automation in production.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search