Site Reliability Engineer, Splunk
Indexed description
This role is a Site Reliability Engineer (SRE) responsible for the architecture, build, and operations of the Splunk logging platform, data ingestion pipelines, and observability capabilities. The focus is on automation and platform-ready delivery, in close collaboration with application, manufacturing, and infrastructure teams, to ensure high-quality ingestion and efficient retrieval of large-scale logs that support troubleshooting and business decisions.
Responsibilities
- Splunk Platform & Reliability
- Design, build, and maintain Splunk infrastructure, including Multi-Site deployments and high-availability (HA) architectures, ensuring reliability, scalability, security, and high performance.
- Operate data-ingestion platforms (pipelines, routing, filtering/enrichment, and integration with Splunk) to improve ingestion quality, cost efficiency, and operability.
- Drive machine-data onboarding, field extraction, search and data-model optimization; create alerts, troubleshoot search performance issues, and define data storage and lifecycle policies.
- Develop automation, monitoring, and diagnostic tools; participate in on-call, respond quickly to bridge calls, minimize incident impact; mentor junior engineers.
- Observability (Log / Metric / Trace)
- Design and implement enterprise observability solutions that support service health assessment, dependency analysis, and fault localization.
- Own architecture, ingestion, storage, query, and capacity management for Prometheus / Mimir; establish dashboard and alerting standards with Grafana.
- Deploy, upgrade, scale, and troubleshoot observability components and collection pipelines in Kubernetes environments.
- Drive correlation and unified views across Log / Metric / Trace; co-define naming, labeling, collection, and SLO/alerting standards.
- Splunk AI OPS
- Design and deliver AI tools and services for troubleshooting, analysis, reporting, and knowledge Q&A (e.g., intelligent search assistance, alert interpretation, root-cause suggestions, runbooks, and ops assistants).
- Build platform AI operations capabilities—alert noise reduction and correlation, intelligent recommendations, and closed-loop feedback—integrating securely with Splunk search, alerts, dashboards, permissions, and audit; measure outcomes via adoption, accuracy, MTTR, and noise rate.
- Document reusable tools and best practices to improve how users leverage Splunk and observability capabilities.
Core Skills
- 7+ years of systems administration experience, with strong Linux background (RHEL / CentOS / Ubuntu, etc.).
- 5+ years designing, maintaining, and troubleshooting mid-to-large-scale Splunk infrastructure; deep understanding of distributed Splunk architecture.
- Strong SPL skills for complex queries; solid grasp of best practices for reports, alerts, and dashboards.
- Hands-on experience operating data-ingestion platforms (collection, processing, routing, and integration with downstream systems).
- Experience with large-scale distributed systems and high-availability architectures; strong analytical and problem-solving skills.
- Experience extending the Splunk ecosystem (Custom Commands, Modular Inputs, App development, REST API, external service integration).
- AIOps experience (alert noise reduction, event correlation, anomaly detection, auto-enrichment, integration with on-call / ticketing / IM).
- Demonstrated ability to use AI effectively (assisted coding, problem analysis, solution design) to continuously deliver maintainable tools, scripts, or automation.
- Configuration management or Infrastructure-as-Code experience (e.g., Ansible, Puppet).
- Experience with OpenTelemetry, APM, or full distributed tracing implementations.
- Experience with long-term metrics storage or federated query platforms (e.g., Mimir / Thanos / Cortex / VictoriaMetrics).
- Experience building or evolving observability platforms on Kubernetes (Operator, Helm, GitOps, etc.).
- Production experience applying LLM / RAG / Agents to operations or data-analysis scenarios.
- Splunk Administrator or Architect level certification.
- Fluent English reading, writing, and speaking.
- Ability to define and drive platform success metrics (onboarding coverage, alert noise rate, MTTR, AI tool adoption, etc.).
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search