Senior Site Reliability Engineer
Indexed description
If you get satisfaction from understanding a production system end to end, finding the real root cause instead of the convenient one, and making the next incident less likely, this role is built for you.
Responsibilities
- Reliability and incident response - Keep the platform available and performant. Define and continuously sharpen monitoring, alerting, and observability. Lead production troubleshooting, drive root-cause analysis, and run post-incident reviews that actually change the system afterwards. Participate in the on-call rotation for the services you own.
- Google Cloud infrastructure and operations: Design, deploy, and optimise our GCP infrastructure - compute, storage, networking, DNS, load balancing, and security services. Drive architecture improvements for reliability, scalability, performance, and cost. Own disaster recovery and business-continuity processes, and prove they work before you need them.
- Containers and orchestration: Build and operate containerised workloads on Docker with a focus on security, performance, and predictable scaling across environments.
- Automation and Infrastructure as Code Provision. Build the scripts and tooling that make operations boring and repeatable.
- CI/CD and release engineering: Maintain CI/CD pipelines and deployment automation. Partner with engineering to make releases safer, faster, and easier to roll back.
- Security and compliance: Apply cloud security practices across IAM, network security, secrets management, and vulnerability remediation. Keep infrastructure aligned to our internal security standards and compliance obligations.
- 5+ years of hands-on Cloud Operations and Site Reliability Engineering, operating production-scale SaaS (not pre-production or internal-only systems).
- You operate Google Cloud at production scale today and can speak in specifics about GCP compute, networking, IAM, GKE, and the operational realities of running real workloads there. This is a hard requirement.
- A second cloud (AWS or Azure) is a plus, not a substitute. We value it, but GCP depth is what the role turns on.
- You debug Linux at the level of "why is this latency spike happening, " not just "restart the service. "
- You reach for automation by reflex. Manual operational work bothers you, and you've built the tooling to remove it.
- Production-scale experience operating cloud infrastructure on Google Cloud.
- Deep Linux systems administration, troubleshooting, and performance tuning.
- Hands-on Docker in production.
- Solid networking fundamentals: VPCs, routing, load balancing, DNS, VPNs, and security controls.
- Monitoring and observability with tools such as Prometheus, Grafana, ELK/OpenSearch, Datadog, or equivalents.
- Scripting and automation in Bash, Python, or similar.
- Git-based workflows and CI/CD pipelines; config management with Ansible or Puppet.
- Strong incident management and root-cause analysis instincts, with a bias toward fixing the system, not the symptom.
- Production experience on a second cloud (AWS or Azure).
- IAM / SSO experience (SAML, OAuth, Okta, or similar).
- Multi-region or multi-cloud operations at scale.
- Background in cloud security, compliance, and governance practices.
- Sustained, high platform uptime against clear SLOs.
- Faster incident detection and resolution, with recurring failure classes systematically driven down.
- More automation and meaningfully less manual operational toil quarter over quarter.
- Observability is good so that the team sees problems before customers do.
- Engineering teams ship reliably because the operational foundation is solid.
- Ownership: You take accountability for outcomes, not just tasks.
- Problem solver: You enjoy diagnosing and resolving complex infrastructure and production challenges.
- Continuous learner: You stay current with evolving cloud, automation, and reliability practices.
- Collaborative: You work effectively across teams and communicate clearly, both in routine operations and in the middle of a critical incident.
- Customer-focused: You understand that infrastructure reliability directly shapes customer experience and business success.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search