Software Engineer - SRE, Retail & Pharmacy
Indexed description
Position Summary
About the Team
Our SRE team is the execution engine behind the reliability, availability, and performance of distributed store technology powering thousands of retail and pharmacy locations nationwide. We operate across pharmacy platforms, Point of Sale (POS) systems, handheld devices, store servers, dispensing systems, and edge computing infrastructure in hybrid cloud and on-premises environments deployed at fleet scale.
Our engineering philosophy is grounded in five pillars: Detection, Prevention, Recovery, Learning Loops, and Developer Experience (DevX).
Our operating principle is what we call the reliability covenant: our success is not measured by how many incidents we respond to, it is measured by how much reliability capability we transfer to the engineering teams we serve. The goal is development teams that carry reliability ownership independently, not teams that rely on SRE to keep their services running. If you are drawn to building capability that outlasts your direct involvement, this team is built for that purpose.
We track operational toil as an engineering metric, not as a permanent operational reality. Engineers are expected to identify recurring manual work, eliminate it through automation, and document the reduction. Toil accumulation is treated as a reliability risk and a capacity cost.
About The Role
As a Software Engineer — SRE, you are a practitioner-level contributor focused on building and running reliable distributed systems. You work within assigned services and domains implementing observability, improving alerting quality, responding to incidents, writing automation, and contributing to the reliability programs that run across the organization.
Scope: Service and task level, you execute with direction and grow toward autonomous ownership.
The environment you are joining
This role exists inside an active SRE transformation. Many of the systems you will monitor, the processes you will contribute to, and the toolchains you will use are being built or significantly improved in parallel with the day-to-day operational work. You will contribute to defining processes as much as following them. Comfort with ambiguity and a bias toward building, not just operating is essential to success in this role. Engineers who thrive here find that environment energizing, not frustrating.
The operating environment includes an edge computing fleet deployed directly inside store locations, unattended nodes that cannot be reached by on-site SRE engineers. This means a deployment or configuration change that goes wrong can simultaneously affect thousands of locations. You will develop a fleet operations mindset alongside a service reliability mindset: blast radius is geographic, not just functional.
What You Will Do
Detection & Observability
- Implement alerting and dashboards using Prometheus, Grafana, Loki, Jaeger, and OpenTelemetry to provide visibility into distributed retail and pharmacy systems
- Contribute to Service Level Indicator (SLI) and Service Level Objective (SLO) definition for assigned services under Senior(s) guidance; learn error budget mechanics and how burn rate translates to patient and customer impact
- Build and maintain service dashboards covering the golden signals: availability, latency, error rate, and saturation anchored to Critical User Journey (CUJ) outcomes, not just infrastructure metrics
- Participate in alert tuning exercises; document false-positive patterns and noise sources with enough specificity to enable SSE-level remediation
- Learn the principles of anomaly-based detection: understand the difference between threshold alerting and time-series baseline deviation, and how ML-generated signals differ from rule-based alerts
- Participate in Production Readiness Reviews (PRR); execute assigned checklist items, contribute findings, and understand the rationale behind each gate
- Write unit and integration tests for SRE tooling and automation scripts at production quality
- Follow and actively improve existing operational runbooks; flag gaps, missing failure modes, outdated steps, and ambiguous procedures for remediation, don't just identify, propose the fix
- Support performance testing and reliability audits under SSE direction; develop hands-on methodology exposure alongside execution skills
- Apply standard reliability engineering patterns: circuit breakers, retries with exponential backoff, timeouts, and bulkhead isolation etc in code and configuration contributions
- Build awareness of fleet-scale deployment risk: understand cohort-based rollout strategies, deployment blast radius concepts, and why configuration drift is a first-order reliability concern in unattended device fleets
- Participate in on-call rotations; escalate promptly per defined escalation paths, speed of escalation is as important as technical accuracy at this level
- Execute validated runbooks during incidents; document all actions, decisions, and timeline entries in real time within incident management systems
- Triage and categorize incoming alerts; route to appropriate domain owners with context, an initial hypothesis, and a timeline of observed signals
- Attend post-incident reviews; capture assigned action items and drive them to closure not just acknowledgment
- Build operational familiarity with the edge computing fleet: Kubernetes clusters at store level, store servers, networking components, and the connectivity failure modes specific to unattended infrastructure
- Attend and actively contribute to postmortems; document timeline details, contributing factors, and observations that would otherwise be lost
- Apply postmortem findings directly: after each postmortem, own at least one monitoring, alerting, or runbook improvement that reduces the likelihood or detection time of the same incident class recurring
- Maintain assigned runbooks and operational documentation in the team's Confluence or wiki space; review for accuracy quarterly and after every related incident
- Complete a structured SRE learning path: SLO fundamentals, incident command, chaos engineering concepts, and retail/pharmacy domain knowledge for the systems you operate
- Track personal MTTD (Mean Time to Detect) and MTTR (Mean Time to Restore) trends across your on-call rotations; use the data to identify your own skill gaps and improvement targets
- Eliminate toil: write automation scripts in Python, Bash, or Go to remove recurring manual operational tasks. For every significant manual task you perform more than twice, your default question should be "why isn't this automated yet?" and your default action should be to fix it
- Maintain a personal record of toil eliminated per quarter, volume of manual touchpoints removed, time recovered, and error modes eliminated
- Contribute to CI/CD pipeline improvements using GitHub Actions, ArgoCD, and Helm; apply GitOps practices for configuration and deployment automation
- Participate in DORA metrics baseline activities; track deployment frequency and lead time for assigned services as indicators of delivery health
- Collaborate with development teams on instrumentation and observability integration during feature development, reliability should be designed in, not added after deployment
- Learn and apply containerization and cloud-native deployment patterns using Kubernetes, Helm, and infrastructure-as-code tools such as Terraform or Ansible
- 2+ years of experience in SRE, DevOps, platform engineering, or related technology roles with production systems responsibility
- 2+ years of experience delivering software in large-scale distributed environments with hands-on application of reliability and resilience concepts
- 1+ year of experience with at least one programming language at production quality: Python, Go, Bash, or Java
- 1+ year of hands-on cloud platform experience: AWS, Microsoft Azure, or Google Cloud Platform (GCP)
- Practical experience with observability and monitoring tools such as Prometheus, Grafana, ELK, Splunk, Datadog, or Dynatrace
- Foundational understanding of containerization and orchestration with Kubernetes and Docker. Experience with AI-assisted tooling and development.
- Comfort contributing to processes that are being defined, not just processes that already exist, ability to operate productively in an environment under active transformation
- Strong written and verbal communication skills; ability to engage both technical and non-technical stakeholders in incident and postmortem contexts
- Experience supporting retail, pharmacy, healthcare, or other distributed-fleet systems at scale
- Exposure to incident management, change management, and problem management processes (ITIL familiarity a plus)
- Familiarity with time-series anomaly detection concepts; understanding of how ML-generated signals differ from rule-based alerting
- Experience with CI/CD pipeline tooling: GitHub Actions, Jenkins, ArgoCD, or CircleCI
- Exposure to microservices architecture, service mesh (Istio, Linkerd), and cloud-native distributed systems patterns
- Basic experience with infrastructure-as-code: Terraform, Ansible, or Pulumi
- Experience with fleet-scale or edge computing deployments where unattended node management is a reliability concern
- You are on-call as a primary/secondary responder, escalating with speed and context, and executing runbooks without hand-holding
- You own and maintain dashboards and SLOs for at least two assigned services, with documented alert quality improvements (fewer false positives, faster detection of real issues)
- You have participated in at least three postmortems and closed every action item assigned to you, with at least one resulting in a monitoring or alerting improvement that reduced recurrence of that incident type
- You have automated at least three recurring manual tasks and documented the before/after reduction in operational toil
- Your peers describe you as someone who closes loops. not someone who documents problems and waits for someone else to fix them
- Bachelor's degree in Computer Science, Engineering, or a related field — or equivalent practical experience
Time Type
Full time
Pay Range
The Typical Pay Range For This Role Is
$72,100.00 - $158,620.00
This pay range represents the base hourly rate or base annual full-time salary for all positions in the job grade within which this position falls. The actual base salary offer will depend on a variety of factors including experience, education, geography and other relevant factors. This position is eligible for a CVS Health bonus, commission or short-term incentive program in addition to the base pay range listed above.
Our people fuel our future. Our teams reflect the customers, patients, members and communities we serve and we are committed to fostering a workplace where every colleague feels valued and that they belong.
Great Benefits For Great People
We take pride in offering a comprehensive and competitive mix of pay and benefits that reflects our commitment to our colleagues and their families.
This full‑time position is eligible for a comprehensive benefits package designed to support the physical, emotional, and financial well‑being of colleagues and their families. The benefits for this position include medical, dental, and vision coverage, paid time off, retirement savings options, wellness programs, and other resources, based on eligibility.
Additional details about available benefits are provided during the application process and on Benefits Moments.
We anticipate the application window for this opening will close on: 10/31/2026
Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state and local laws.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search