Data Platform Observability Engineer
Indexed description
Data Platform Observability Engineer
Location: Charleston, SC
Clearance Level: Active Secret Security Clearance Required
Direct-Hire & Full-Time
Hybrid: Schedule 2 days remote, 3 days in office per week
We are seeking an experienced Data Platform Observability Engineer to help design, implement, and promote modern observability capabilities supporting secure, mission-critical systems.
This position will work within a broader data-engineering organization but will have independent ownership of the observability platform. The engineer will help technical teams collect, route, visualize, and interpret application and platform telemetry—including metrics, logs, and distributed traces.
The strongest candidates will understand the complete observability process: instrumenting services, collecting telemetry, moving it through the platform, storing it in the appropriate backend, building meaningful dashboards and alerts, and using the resulting information to troubleshoot complex systems.
Required Qualifications:
- 3+ years of hands-on experience with observability, monitoring, site reliability, platform engineering, DevOps, or a closely related discipline.
- Practical experience working with application and infrastructure metrics, logs, traces, dashboards, and alerts.
- Hands-on knowledge of Prometheus or a comparable metrics platform.
- Experience creating operational dashboards with Grafana, Kibana, Splunk, Dynatrace, or similar visualization tools.
- Experience with at least one modern logging or tracing architecture, such as Loki, Tempo, Elastic, Splunk, or Dynatrace.
- Ability to explain how telemetry is generated, collected, processed, stored, queried, visualized, and used during troubleshooting.
- Familiarity with Docker, Kubernetes, and containerized application environments.
- Understanding of Agile delivery practices and collaborative software-development processes.
- Strong troubleshooting, documentation, communication, and problem-solving skills.
- Ability to work independently while coordinating effectively across multiple technical teams.
Preferred Qualifications:
- Experience implementing or migrating services to OpenTelemetry.
- Experience with OpenTelemetry SDKs, Collectors, receivers, processors, exporters, and context propagation.
- Familiarity with the Grafana LGTM stack: Loki, Grafana, Tempo, and Mimir.
- Experience streaming or routing telemetry through Kafka, Cribl, or comparable technologies.
- Experience deploying applications through Kubernetes and Helm.
- Familiarity with GitLab CI/CD, GitHub Actions, Azure DevOps, or similar delivery pipelines.
- Experience with Terraform, Ansible, Chef, Puppet, or another automation platform.
- Programming or scripting experience with Python, PowerShell, Bash, Go, or Java.
- Familiarity with Nginx, Kubernetes ingress, API routing, reverse proxies, and load balancing.
- Experience supporting data platforms, microservices, real-time pipelines, or distributed systems.
- Bachelor’s degree in computer science, engineering, information technology, or a related field; equivalent professional experience may be considered.
Essential Duties:
- Design, implement, and maintain observability capabilities for data platforms, applications, pipelines, and distributed services.
- Collect and correlate metrics, logs, and traces to provide a unified view of platform health and performance.
- Develop dashboards, alerts, service-level indicators, service-level objectives, and operational reports.
- Build visualizations using Grafana, Kibana, Splunk, Dynatrace, or comparable platforms.
- Configure and support Prometheus for metrics collection, querying, alerting, and performance monitoring.
- Work with supporting observability technologies such as Loki for logs and Tempo for distributed traces.
- Implement and maintain OpenTelemetry instrumentation, collectors, pipelines, processors, and exporters.
- Support the modernization of legacy monitoring and logging solutions through adoption of OpenTelemetry standards.
- Monitor service availability, latency, throughput, errors, resource consumption, storage utilization, and data-pipeline failures.
- Troubleshoot application, data-flow, infrastructure, Kubernetes, and network issues using correlated telemetry.
- Route, buffer, transform, and deliver observability data using technologies such as Kafka, Cribl, or similar platforms.
- Containerize and deploy observability services using Docker, Kubernetes, and Helm.
- Support ingress, reverse-proxy, load-balancing, and service-routing configurations using Nginx or comparable technologies.
- Automate platform deployment, configuration, testing, and operational procedures.
- Participate in architecture reviews, technical design discussions, peer reviews, and Agile sprint activities.
- Create architecture diagrams, implementation guides, operating procedures, and onboarding documentation.
- Partner with software, platform, infrastructure, cybersecurity, and data-engineering teams to understand their telemetry requirements.
- Present observability capabilities to other engineering teams and help them adopt the platform, standards, and services.
Education:
- Bachelor’s degree in computer science, engineering or a related field preferred; equivalent relevant professional experience may be considered.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search