Back to search
e&e IT Consulting Services, Inc. Linkedin · Posted 13d ago

Site Reliability Engineer

Harrisburg

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

e&e is seeking a Site Reliability Engineer an onsite contract opportunity in Harrisburg, PA!


We are seeking an experienced Site Reliability Engineer (SRE) to support the reliability, availability, and operation of mission-critical on-premises and self-hosted systems. This role is focused on production operations, observability, automation, incident response, and disaster recovery within a hands-on infrastructure environment.


The Site Reliability Engineer will own systems end to end, helping ensure that critical services remain available, observable, scalable, and recoverable. The ideal candidate brings a reliability-first mindset, strong Linux and Kubernetes expertise, and experience reducing operational toil through automation. This position will also partner closely with development teams to improve production readiness, capacity planning, change management, and overall platform resilience.


Key Responsibilities

  • Maintain the availability, reliability, and overall health of on-premises and self-hosted production systems.
  • Define, measure, monitor, and maintain Service Level Objectives (SLOs) for critical services.
  • Build and maintain comprehensive observability across platforms using metrics, logs, traces, dashboards, and alerting.
  • Participate in and support on-call operations, incident detection, troubleshooting, mitigation, and service restoration.
  • Lead or contribute to post-incident reviews and postmortems, identifying root causes and implementing durable corrective actions.
  • Automate provisioning, deployment, recovery, maintenance, and other operational processes to reduce manual effort and operational toil.
  • Operate and troubleshoot core platform technologies, including Kubernetes, Linux, databases, networking, and secrets management.
  • Improve disaster recovery capabilities and develop reliable, repeatable processes for datacenter and service restoration.
  • Support Kubernetes deployments, networking, configuration, upgrades, and production troubleshooting.
  • Develop scripts and automation using technologies such as Python, Bash, or Go.
  • Implement and maintain infrastructure-as-code and deployment automation using Helm, GitOps, and related technologies.
  • Support CI/CD processes and integrations, including GitLab-based pipelines.
  • Partner with application and development teams to assess production readiness, capacity requirements, system dependencies, and operational risks.
  • Participate in incident, problem, and change management processes.
  • Continuously identify opportunities to improve system resilience, recoverability, performance, and operational efficiency.


Requirements

  • Strong experience operating production environments in a Site Reliability Engineering, Platform Engineering, DevOps, Infrastructure, or Operations Engineering capacity.
  • Demonstrated reliability-first mindset with a focus on availability, resilience, automation, and operational excellence.
  • Strong Linux administration and troubleshooting skills, including the ability to diagnose complex production issues under pressure.
  • Hands-on experience operating Kubernetes in production, including deployments, networking, configuration, and failure troubleshooting.
  • Experience implementing and supporting observability solutions across metrics, logging, and distributed tracing.
  • Knowledge of defining, measuring, and managing SLOs and related reliability metrics.
  • Experience with monitoring and observability technologies such as Prometheus, Grafana, or comparable platforms.
  • Strong automation and scripting skills using Python, Bash, Go, or similar technologies.
  • Experience with infrastructure-as-code, Helm, GitOps, and automated deployment practices.
  • Hands-on incident response and on-call experience, including root cause analysis and postmortem development.
  • Experience working with APIs, systems integrations, and CI/CD pipelines.
  • Experience supporting self-hosted, on-premises, bare-metal, or virtualized infrastructure is highly preferred.
  • Experience with Cilium, eBPF networking, service mesh technologies, ClusterMesh, or cross-cluster service discovery is a plus.
  • PostgreSQL operational experience, including replication, failover, and database operators such as Crunchy, is preferred.
  • Experience with secrets management platforms such as Vault or OpenBao is a plus.
  • Knowledge of disaster recovery and multi-datacenter reliability architecture is preferred.
  • Experience with identity-aware or zero-trust access technologies such as Teleport is beneficial.
  • Familiarity with ITSM/ITIL practices, including incident, problem, and change management.
  • Experience operating systems subject to regulatory, security, or data-privacy requirements is preferred; familiarity with FERPA or similar standards is a plus.


Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search