SRE Observability SLO Engineer
Indexed description
Job Description
Roles and Responsibilities
Telemetry Standards & Architecture
- Implement organization-wide telemetry standards covering metrics, logs, and distributed traces across all GridOS SaaS services.
- Implement metrics collection for Kubernetes-hosted services (EKS/Rancher) including pod-level, namespace-level, and cluster-level metrics.
- Working with the SRE Lead and SRE Platform Engineers help define and implement data retention policies, cardinality budgets, and telemetry cost controls to keep observability economically sustainable.
- Publish and maintain an Observability Runbook library covering onboarding, alert tuning, and dashboard standards for Platform SRE and Production DevOps teams.
- Partner with product engineering, Platform SRE, and customer stakeholders to define meaningful Service Level Indicators (SLIs) and Service Level Objectives (SLOs) per product and customer tier.
- Build and maintain SLO tooling — error budget burn-rate alerts, burn-rate dashboards, and automated SLO compliance reports.
- Govern the SLO review cycle: facilitate monthly SLO reviews, identify reliability risks early, and drive prioritization of reliability work with the SRE Lead.
- Translate SLOs into SLAs for customer-facing commitments in coordination with the SRE Team Lead.
- Design and build operational dashboards covering availability, latency, error rates, and saturation (the 'Golden Signals') for every GridOS SaaS product.
- Implement alert policies with noise-reduction practices: symptom-based alerting, multi-window burn-rate rules, and alert deduplication.
- Create executive-level dashboards for SRE leadership and customer-facing uptime/availability reports aligned to contractual SLAs.
- Establish and maintain alert routing, escalation policies, and on-call schedules in coordination with the incident response workflow.
- Design and implement a synthetic monitoring plan covering critical user journeys for each GridOS SaaS product and customer environment.
- Build synthetic checks for API health, UI flows, and integration endpoints using AWS CloudWatch Synthetics or equivalent tooling.
- Define alerting thresholds for synthetic monitors and integrate them into the broader incident detection pipeline.
- After v1.0 delivery, transition into a roadmap-aligned improvement cycle: expand coverage for new features, tune alert signal-to-noise, and retire stale monitors.
- Conduct periodic observability health reviews to identify gaps in coverage, reduce MTTD (Mean Time to Detect), and improve MTTR (Mean Time to Resolve).
- Collaborate with the Production DevOps engineer on FinOps validation — correlate infrastructure cost metrics with performance and reliability data.
- 2–3 years in SRE, observability engineering, or infrastructure reliability roles.
- Deep expertise with at least one major observability platform — Datadog, Grafana + Prometheus, AWS CloudWatch, Dynatrace, or New Relic.
- Hands-on experience implementing SLIs, SLOs, and error budget burn-rate alerting in a production SaaS environment.
- Strong understanding of distributed systems telemetry: metrics (Prometheus/CloudWatch), structured logging (CloudWatch Logs Insights, ELK), and distributed tracing (OpenTelemetry, AWS X-Ray).
- Experience with Kubernetes observability — kube-state-metrics, node exporters, Helm-deployed monitoring stacks, and namespace-level resource metrics.
- Proficiency in at least one query/visualization language: PromQL, Splunk SPL, Datadog Query Language, or CloudWatch Logs Insights query syntax.
- Experience designing alerting strategies that minimize alert fatigue through symptom-based and burn-rate approaches.
- Scripting skills in Python and/or Bash for automation of monitoring configuration and report generation.
- Cloud Technologies - AWS Cloud Infrastructure - EKS, RDS, MSK, S3, EC2, EBS, SQS, etc.
- Kubernetes - EKS, Rancher
- Infrastructure as Code: Terraform
- Deployment and Configuration Tools - Ansible, Chef or Puppet
- Telemetry standards and tools - Open Telemetry, CloudWatch, Cloudtrail
- Observability tools and technology - Datadog, Splunk, NewRelic, etc.
- Alerting and notification - AWS and Azure alerting notification
- Scripting - Go, Python, Groovy, Bash
- Strong Linux Administration Skills
- Strong analytical and problem solving skills
- Familiarity with OpenTelemetry (OTel) for vendor-agnostic instrumentation.
- Experience with synthetic monitoring tools — AWS CloudWatch Synthetics, Datadog Synthetics, or Catchpoint.
- Knowledge of chaos engineering practices for reliability validation (Chaos Monkey, AWS Fault Injection Simulator).
- Exposure to AIOps or ML-driven anomaly detection features within observability platforms.
- Experience in regulated industries — energy, utilities, healthcare — where compliance-grade audit trails are required.
- AWS certifications: CloudWatch / Observability specialty, Solutions Architect Associate or Professional.
Leadership
- Influences through others; builds direct and "behind the scenes" support for ideas.
- Preemptively sees downstream consequences and effectively tailors influencing strategy to support a positive outcome.
- Able to verbalize what is behind decisions and downstream implications.
- Continuously reflecting on success and failures to improve performance and decision-making.
- Understands and encourages change when needed.
- Proactively identifies and removes project obstacles or barriers on behalf of the team.
- Able to navigate accountability in a matrixed organization.
- Self-starter; communicates and demonstrates a shared sense of purpose. Learns from failure.
- Critical thinker; able to quickly adapt to changing environments
- A hacker or tinkerer at heart
- Risk taker, not afraid to think outside the box or challenge the status quo
- Emotional Intelligence, ability to influence up and out and the ability to work independently
- Must be a team player with a strong desire to win
- Passionate about continuously learning
- Highly organized and efficient; able to balance competing priorities and execute accordingly
- Strong oral and written communication skills.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search