Back to search
eTeam Linkedin · Posted 22d ago

Observability Architect

Troy, New York, United States

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Job Title: Observability Architect

Location: Troy NY

Pay Range: $80 - 87/hr on W2

Duration: 12 Months (Possibility of Extension)


Summary

Senior Observability Architect responsible for platform modernization, Grafana Cloud migration, operational excellence, production support readiness, and enterprise observability strategy across hybrid cloud environments.


Core Responsibilities

  • Lead observability strategy, architecture, and operational excellence initiatives.
  • Provide servant leadership to Operations and Production Support teams.
  • Own end-to-end design across all layers (infrastructure → end user); define standards, SLOs, and signal model.
  • Drive continuous improvement programs focused on reliability, system health, and MTTR reduction.
  • Ensure operational readiness and supportability of all platform changes before production deployment.


Observability Platform Modernization

  • Mature the enterprise observability platform and lead migration from New Relic to Grafana Cloud at scale.
  • Design SLO/SLI frameworks, anomaly detection (Sift), and predictive alerting capabilities.
  • Design observability standards for AWS and on-premises hybrid environments.
  • Standardize telemetry collection using Grafana Alloy and OpenTelemetry.
  • Develop enterprise monitoring, logging, tracing, dashboarding, and alerting frameworks.


Grafana & OpenTelemetry Architecture

  • Architect solutions leveraging Grafana Mimir, Loki, Tempo, IRM, and Grafana Cloud.
  • Define instrumentation patterns and validated templates for infrastructure, applications, and distributed services.
  • Design dashboards, alerts, service health views, SLIs, SLOs, and error-budget monitoring.
  • Ensure telemetry quality, scalability, and governance across the organization.


Production Support & Operational Readiness

  • Own production support intake, triage, escalation, and stakeholder communications.
  • Validate monitoring, alerting, logging, and support coverage prior to go-live.
  • Coordinate incident response, root cause analysis, and service restoration activities.
  • Improve Mean Time to Detect (MTTD) and Mean Time to Resolution (MTTR).


Runbooks & Knowledge Management

  • Create, maintain, and continuously improve Tier 2 and Tier 3 runbooks.
  • Develop onboarding documentation, templates, and operational standards.
  • Partner with support teams to ensure consistent observability adoption and execution.


Tools & Technologies

  • Grafana Cloud
  • Grafana Alloy
  • OpenTelemetry
  • Grafana Mimir
  • Grafana Loki
  • Grafana Tempo
  • Grafana IRM
  • Splunk
  • New Relic
  • Power BI
  • SQL
  • AWS Cloud Services
  • Hybrid Infrastructure Monitoring


Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search