Back to search
MX Linkedin · Posted yesterday

Senior Observability Engineer

India

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

We are fueled by a moral imperative to advance mankind, and it all begins with our people, our product, and our purpose. Passion isn’t something we turn on and off; it’s woven into everything we do. If you thrive in high-challenge environments, are inspired by exceptional teammates, and are driven to grow beyond what you thought possible, MX is where you belong.

Come build the future with us. Join an award-winning company that isn’t just shaping the financial industry, but transforming it in ways that create meaningful, lasting impact for millions of people.

At MX, reliability is a product. Our infrastructure powers financial applications used by millions of people and processes billions of transactions for major financial institutions, and customers feel every second of downtime.

We're building a new observability function that runs the way we run incident response: the system does the heavy lifting, and people handle judgment, customers, and the exceptions. As a Senior Observability Engineer, you build and operate an observability control plane. You scaffold baselines, score coverage, and turn every real incident into the detection the platform should have caught. This is a multiplier role: you raise the bar for every team through standards and automation instead of building each team's dashboards by hand.

We call it the shepherd model. You shepherd Datadog and partner with our product engineering teams so they observe the right signals for their products. Service owners get real signal instead of noise, and leadership gets coverage and health as a program metric.

This role shares the team pager. Observability and incident response run one on-call roster. You take shifts with the rest of the team and act as Incident Commander when an incident needs one. It is core to the role, not an afterthought.

Engineering at MX runs hybrid infrastructure (AWS and bare metal) with services in Ruby, Go, and Java, messaging over NATS and RabbitMQ, and data on PostgreSQL and Redis. Datadog is our observability platform and incident.io is our incident response platform.

Job Duties

  • Build and operate an observability control plane: automate baseline monitors, dashboards, and tagging standards through the Datadog API and Terraform.
  • After significant incidents, produce detection and dashboard gap packs grounded in Datadog and MX investigation patterns, with queries ready to apply.
  • Define what "good" looks like for a Ruby, Go, or Java service on Datadog (tags, golden signals, alert quality, dashboard contracts), then audit services against that standard and accept or reject readiness.
  • Validate, don't own. Service owners keep their alerts and dashboards; you confirm they are complete and correct, then move on. Escalate to engineering managers when coverage fails or an owner is missing.
  • Own the monthly observability and service-catalog health report: departed owners, stale dashboards, services with no monitors, SLO gaps, and coverage trends.
  • Run maturity assessments (baseline through SLO, launch-ready, self-serve) and track them over time.
  • Tune alerting toward zero false SEV1/2 pages and actionable SEV3/4 alerts, and coach teams on Datadog cost and cardinality.
  • Build self-serve onboarding so new services get baseline observability on day one, without a multi-week embed.
  • Share the team pager. Rotate on the shared IR & Observability on-call, triage and investigate live incidents with Datadog and MX investigation patterns, and take Incident Commander or supporting technical roles as the incident needs.
  • After incidents, close the detection loop (gap packs, new monitors, dashboards) so the pager gets quieter over time.
  • Run high-value launch and production-readiness reviews as a checkpoint, not a permanent staffing model.

Basic Requirements

  • BS in Computer Science or equivalent experience
  • 5+ years running production observability, SRE, or DevOps. Datadog preferred; strong Grafana/Prometheus, Splunk, or New Relic experience counts if you can ramp on Datadog fast.
  • Automation-first engineering in Python, Bash, Go, and/or Terraform, plus Kubernetes proficiency. You encode monitoring standards as code rather than clicking the UI.
  • AI- and workflow-literate. You've used or built scripted and AI-assisted workflows to scale reviews, audits, and docs.
  • Alerting and SLO strategy: burn-rate and error-budget thinking, with a track record of cutting alert fatigue on evidence.
  • Distributed-systems debugging across microservices: latency, connection pools, queues, and cascading failure on Kubernetes and bare metal, with NATS, RabbitMQ, Postgres, and Redis.
  • Shared on-call, Incident Commander-capable. You've run or supported incident bridges and written postmortems, and you'll take shifts on the shared IR & Observability rotation.

Preferred Requirements

  • Fintech experience with MX-like architectures
  • Google SRE practices: toil elimination, incident management, automation for self-healing
  • Cross-functional influence without authority. You've improved teams that don't report to you.
  • Governance and reporting: you can produce a monthly health and compliance report leadership reads (orphans, stale entries, gaps, trends).
  • OpenTelemetry instrumentation
  • Datadog cost optimization at scale (cardinality, log indexing, sampling)
  • Incident response platforms (incident.io, PagerDuty, OpsGenie); prior formal Incident Commander experience
  • Golang and Ruby on Rails (the MX stack)

What Success Looks Like

By six months, you're a full participant on the shared on-call rotation and a capable Incident Commander on live SEVs, teams you've engaged have alerts and dashboards that answer "what's broken and where do I look?", and the monthly health report runs largely on its own. By twelve months, incidents get caught earlier because of instrumentation the loop added, new services get baseline observability from a self-serve template on day one, and no team depends on a shepherd for day-one coverage.

Compensation

The expected earnings for this role could be comprised of a base salary and other forms of cash compensation, such as bonus or commissions as applicable.

This pay range is just one component of MX’s total rewards package. MX takes a number of factors into account when determining individual starting pay, including job and level they are hired into, location, skillset, peer compensation.

  • Please note applicants applying for this position must have the legal right to work in India without the need for sponsorship. We are unable to provide work sponsorship for this role, and candidates should be able to verify their eligibility to work in the country independently. Proof of eligibility to work in India will be required as part of the hiring process.

Work Environment

In this role, a significant aspect of the job involves working in the office for a standard 40-hour workweek. We believe that the collaborative nature of our work and the face-to-face interactions among team members are essential for fostering a dynamic and productive work environment. Being present in the office enables seamless communication, facilitates quick decision-making, and encourages spontaneous collaboration that contributes to the overall success of our projects. We value the synergy that comes from having our team members physically together, allowing for immediate problem-solving, idea exchange, and team building.

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search