Back to search
Mumba Technologies, Inc. Linkedin · Posted 10d ago

Site Reliability Engineer

India

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Job Title: Site Reliability Engineer

Job Type: Full Time

Location: Gurgaon (Hybrid 3-Days in the office)


About the Role:

We are looking for an experienced Site Reliability Engineer (SRE) to help build and operate highly reliable, scalable, secure, and observable systems. The role involves hands-on work across AWS, Kubernetes, Infrastructure as Code, observability, automation, security, and AI-driven SRE practices.


Reliability & Availability

  • Lead incident response, conduct RCAs and ensure action items are tracked to closure
  • Build and maintain runbooks, playbooks and escalation frameworks for proactive and reactive response
  • Drive toil reduction by identifying repetitive operational work and engineering it away


Observability

  • Design and own the full observability stack — metrics, logs, traces and events — using tools like Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace or similar
  • Build intelligent alerting that reduces noise, eliminates alert fatigue and surfaces actionable signals
  • Implement distributed tracing and dependency mapping to provide end-to-end visibility across microservices
  • Drive adoption of continuous profiling and real user monitoring (RUM) for proactive performance management


AI Adoption in SRE

  • Leverage AIOps platforms to enable anomaly detection, predictive alerting and automated root cause analysis
  • Implement AI-assisted incident triage — using LLM-powered tools to summarise incidents, suggest fixes and accelerate MTTR
  • Build and maintain ML-powered capacity forecasting models to optimise infrastructure spend and prevent resource saturation


Security & Vulnerability Management

  • Embed security-as-reliability principles — treating security incidents with the same urgency as availability incidents
  • Own issue remediations across infrastructure (OS, containers, dependencies)
  • Integrate SAST, DAST and SCA tools into CI/CD pipelines to shift security left


Infrastructure & Platform Engineering

  • Design, build and maintain cloud-native infrastructure on AWS using Infrastructure as Code (Terraform, Pulumi) & drive rightsizing, reserved capacity planning and cost anomaly detection
  • Own Kubernetes cluster operations — autoscaling, resource management, networking and upgrade strategy


Leadership & Culture

  • Mentor and guide junior and mid-level SREs — conducting technical reviews and pair debugging sessions
  • Define and evolve SRE team standards, best practices and engineering principles
  • Collaborate closely with product, development and security teams as an embedded reliability partner
  • Contribute to on-call rotation and drive continuous improvement of on-call experience
  • Represent SRE in architecture reviews, sprint planning and cross-functional forums


The Competitive Edge

AEM Administration

  • Own end-to-end reliability and availability of AEM environments — Author, Publish, Dispatcher and AEM as a Cloud Service (AEMaaCS) — across dev, staging and production
  • Monitor and manage AEM instance health, optimise Dispatchers, Manage DAM, OSGi Configurations, replication queues.


Exposure to CDN

  • Experience with Cloudflare - CDN, Workers.


Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search