L2 Support Engineer - Managed Services
Indexed description
The Level 2 Managed Service Engineer for the End-to-End Observability team is responsible for advanced troubleshooting, root cause analysis, Datadog configuration, incident escalation, and continual service improvement across the observability stack. This role serves as the primary escalation point for L1 engineers and is expected to independently resolve, or drive resolution of, moderately complex incidents spanning application, infrastructure, network, APM, Synthetics, logs, and RUM data across both 3-tier and cloud-native environments.
Key Responsibilities
- Act as the first escalation point for incidents triaged by L1 engineers that require deeper investigation.
- Perform root cause analysis (RCA) by correlating APM traces, infrastructure metrics, network telemetry, logs, Synthetics results, and RUM data.
- Configure and fine-tune advanced Datadog capabilities: custom metrics, log pipelines, monitors with composite/anomaly detection, SLOs, and dashboards aligned to service ownership.
- Guide and review Datadog APM instrumentation coverage across services; identify monitoring gaps and recommend improvements.
- Independently identify whether an issue originates in the application tier, infrastructure, network, or a third-party dependency, in both 3-tier and cloud-native (microservices/containerized) environments.
- Lead or actively participate in major incident bridge calls, providing observability data to speed up resolution.
- Own problem management: identify recurring incidents, drive permanent fixes, and track them to closure.
- Mentor and coach L1 engineers on triage techniques, Datadog usage, and escalation quality.
- Contribute to automation of routine monitoring tasks (alert enrichment, auto-remediation triggers, ticket auto-creation).
- Support capacity and performance trend analysis using historical Datadog data.
- Review and refine monitors to reduce noise and false positives (alert tuning).
- Ensure adherence to SLAs/OLAs and maintain quality of incident documentation and RCA reports.
- Participate in client/stakeholder calls to explain incidents, trends, and improvement plans.
- 3 to 6 years of experience in Managed Services, NOC, SRE, or application/infrastructure support roles, with at least 1-2 years working hands-on with Datadog.
- Strong working knowledge of Datadog across monitoring, dashboards, alerting, APM, Synthetics, RUM, logs, metrics, SLOs, and service-level troubleshooting.
- Advanced monitor configuration (anomaly detection, forecast, composite monitors, SLOs).
- APM: distributed tracing, service maps, trace analytics, and identifying instrumentation gaps.
- Synthetics: designing multi-step API and browser tests, analyzing failure trends.
- RUM: correlating front-end performance and errors with backend traces.
- Log Management: building log pipelines, parsing rules, and log-based metrics.
- Solid understanding of infrastructure concepts: compute, storage, containers, orchestration (Kubernetes basics), and cloud resource scaling.
- Solid understanding of networking fundamentals: DNS, load balancing, latency/packet loss analysis, firewall and connectivity troubleshooting.
- Strong analytical ability to read application logs, stack traces, and error patterns to isolate the failing component, without needing to be an expert in the underlying code.
- Experience with ITSM tools for incident, problem, and change management.
- Working knowledge of ITIL processes, especially incident and problem management.
- Solid understanding of 3-tier application architecture and how load, latency, or failure at one tier cascades to others.
- Solid understanding of cloud-native architecture: microservices, containers, API gateways, service meshes, and inter-service communication patterns.
- Ability to correlate application-level symptoms (errors, latency, timeouts) with underlying infrastructure or network causes using observability data alone.
- Understanding of common resiliency patterns (retries, circuit breakers, caching, load balancing) and how their failure shows up in monitoring data.
- Datadog certification (APM, Infrastructure, or Log Management specialization).
- Scripting experience (Python, Shell, PowerShell) for automation and custom integrations.
- Exposure to Infrastructure as Code (Terraform) for managing Datadog configuration as code.
- Experience with CI/CD pipelines and how deployments correlate with monitoring anomalies.
- Exposure to other observability or APM tools (New Relic, Dynatrace, Splunk, Grafana, Prometheus).
- Strong ownership and accountability for incident resolution.
- Ability to communicate technical findings clearly to both technical teams and stakeholders.
- Mentoring ability and patience when guiding L1 engineers.
- Comfortable leading discussions during high-pressure major incidents.
- Structured approach to documentation and knowledge sharing.
- Bachelor's degree in Computer Science, Information Technology, Electronics, or a related field (or equivalent practical experience).
- Level: L2 (Mid-level, Managed Services)
- Shift: 24x7 rotational shifts, with on-call/escalation coverage as needed
- Reporting To: L3 Engineer / Observability Team Lead
Skills: rum,infrastructure,shell,apm,datadog,itil,dns,powershell scripting,python,container orchestration,splunk
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search