Act as software detectives, provide a dynamic service identifying and solving issues within multiple components of critical business systems.
Deep Technical Support & Root Cause Analysis
Handle complex customer issues, going deep into root cause analysis that spans customer configuration, product bugs, and underlying infrastructure.Proactive Reliability Improvement
Analyse patterns of customer issues, support cases, and outage impacts to identify systemic weaknesses.
Design and implement solutions, automation, or monitoring to prevent future occurrences (involving coding, configuration changes, or proposing architectural improvements).Bridging Customer Impact and Engineering
Act as a liaison between the customer-facing support teams and the core Support/Development teams.
Translate customer pain into technical requirements and SLOs, and explain technical constraints and incident impacts back to support teams.Improving Supportability
Develop tools, playbooks, and dashboards to help front-line support and diagnose and resolve issues more quickly.
Feedback into the product development lifecycle to ensure new features are designed with supportability and reliability in mind.Customer-Centric Service Level Objectives (SLOs)
Contribute to defining and refining SLOs to better reflect actual customer pain and perceived performance, not just server-side metrics.Incident Management & Postmortems
Participate in incident response, bringing a strong understanding of customer impact.
Contribute significantly to postmortems, ensuring preventative actions address both the technical root cause and the customer's experience.Proactive Customer Engagement (for Key Customers)
Engage with key customers to understand their critical workloads, review their architecture, and provide guidance on reliability best practices.EssentialExperience & Qualifications
4–8+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Systems Engineering.
Experience supporting enterprise platforms.
Experience with cloud-native technologies and automation.
Experience with AI / MLDesirable
Certified Kubernetes Administrator (CKA) / Certified Kubernetes Application Developer (CKAD)
Google Professional Cloud DevOps Engineer
Google Professional Cloud Network Engineer / VMware Certified Professional - Network Virtualization (2V0-41.24)
Linux Foundation Certified System Administrator (LFCS) / Linux Foundation Certified IT Associate (LFCA) / Linux - ------- Professional Institute LPIC-3 Mixed Environments
Prometheus Certified Associate (PCA)Required Technical SkillsInfrastructure & Platforms
Kubernetes (essential)
Linux (high priority)
Networking (high priority)
Service Mesh Basics (Istio) (high priority)
PKI (high priority)
AI / ML (high priority)
Google Compute
Storage Technologies (optional)
Observability & Monitoring
Prometheus (high priority)
Grafana (high priority)
Loki (high priority)
Splunk