Observability Technical Lead
Indexed description
Key Responsibilities
Enterprise Observability Platform Engineering
- Serve as the senior SME for enterprise observability, defining architectures, standards, integration patterns, and reusable solutions
- Design and optimize observability across cloud, on-prem infrastructure, networks, Kubernetes, containers, applications, APIs, databases, middleware, and enterprise platforms
- Standardize metrics, logs, traces, events, topology, dashboards, alerting, instrumentation, and service health
- Provide technical leadership for LogicMonitor, Splunk, OpenTelemetry, Grafana, Prometheus, Tempo, VictoriaMetrics, Loki, and related technologies
- Lead OpenTelemetry adoption, including instrumentation, collectors, telemetry pipelines, distributed tracing, context propagation, and vendor-neutral standards
- Design telemetry pipelines that route metrics, logs, and traces across multiple platforms
- Develop API-driven and Observability-as-Code capabilities for onboarding, configuration, validation, and lifecycle management
- Establish governance for RBAC, telemetry standards, alerting, retention, integrations, configuration management, and data lifecycle
- Improve observability coverage, telemetry quality, scalability, reliability, performance, and cost efficiency
- Drive automation and self-service through APIs, Infrastructure-as-Code, CI/CD, GitOps, and reusable observability pattern
- Integrate Edwin AI and AIOps capabilities for anomaly detection, event correlation, root-cause analysis, investigation, and operational intelligence
- Connect metrics, logs, traces, events, topology, CMDB, service ownership, and operational context to deliver end-to-end visibility
- Establish reliability standards including SLIs, SLOs, error budgets, monitoring coverage, alert quality, MTTD, and MTTR
- Reduce alert fatigue through intelligent correlation, dynamic thresholds, suppression, automation, and event-management practices
- Evaluate eBPF, continuous profiling, Kubernetes observability, dependency mapping, and auto-instrumentation
- Advance operations from reactive monitoring toward proactive and predictive operations
- Collaborate with engineering and operations teams across North America, Europe, and Asia
- Lead architecture reviews, platform evaluations, workshops, and technical working sessions
- Influence observability strategy and standards across teams without direct authority
- Mentor engineers and promote observability, SRE, and reliability best practices
- Evaluate emerging technologies and recommend adoption based on interoperability, scalability, business value, and cost
- Translate strategy into standards, reference architectures, and reusable implementation patterns
- Bachelor’s Degree
- 7+ years of experience in observability, monitoring, APM, SRE, DevOps, platform engineering, or related professional experience
- 7+ years of experience designing, implementing, operating, and maintaining enterprise-scale observability platforms using observability-as-code, monitoring-as-code, infrastructure-as-code, GitOps, and CI/CD practice
- Experience with AWS, Azure, GCP, and large-scale Kubernetes environments
- Experience designing OpenTelemetry Collector architectures and telemetry pipelines
- Expertise with Grafana, Tempo, Loki, and Prometheus-compatible platforms
- Experience with VictoriaMetrics or other large-scale time-series databases.
- Knowledge of eBPF, continuous profiling, auto-instrumentation, and cloud-native telemetry
- Experience integrating observability platforms with ServiceNow, ITSM, CMDB, incident management, and automation platforms
- Understanding of SRE practices including SLIs, SLOs, error budgets, and incident management
- Experience enabling developer self-service and internal developer platform integrations
- Knowledge of RBAC, secrets management, governance, compliance, and telemetry data protection
Other Compensation
This position is entitled to short-term cash incentives, subject to plan requirements.
Benefits
Employees are eligible for benefits, including:
- Health Care Benefits: Medical, Dental, Vision; Wellness incentives
- Retirement Benefits
- Time off and Leave: Paid vacation days, up to 15 days; paid sick days, up to 5 days; paid personal leave, up to 5 days; paid holidays, up to 13 days; birth and adoption leave; parental leave; family and medical leave; bereavement leave; jury duty leave; military leave; purchased vacation
- Disability: Short-term and long-term disability
- Life Insurance and Accidental Death and Dismemberment
- Tax-Advantaged Accounts: Health Savings Account; Health Care Spending Account; Dependent Care Spending Account
- Tuition Assistance
Carrier EEO Statement and Accommodations Process
Carrier is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability or veteran status or any other applicable state or federal protected class. Carrier provides affirmative action in employment for qualified individuals with a Disability and Protected Veterans in compliance with section 503 of Rehabilitation Act and the Vietnam Era Veterans' Readjustment Assistance Act.
If you require a reasonable accommodation to complete the application process, participate in an interview, or otherwise engage in the hiring process, please contact us at [email protected] (opens in new window). We will make every effort to meet your needs in accordance with applicable laws.
Application Deadline
Applications will be accepted for at least 3 days from Job Posting Date: 9 September 2026
Job Applicant's Privacy Notice
Please click on the link to review the Job Applicant Privacy Notice (opens in new window).
Use of AI
Technology-enabled tools may support parts of the recruitment process, with oversight by people.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search