Senior Software Engineer, DevOps
Indexed description
Our mission is to make healthspan and lifespan equal for all by translating science into medicine in real-time, all while bringing humanity back into health care. Delivering such robust, personalized, and preventive health care is complex and requires a team-wide dedication to excellence.
After successfully opening our flagship Institute in New York in 2022 and expanding to South Florida in 2024, we are now bringing the Atria experience to the West Coast with the launch of our Los Angeles Institute in late spring 2026.
Atria Health is seeking a Senior Software Engineer for our DevOps team to help design, build, and operate the infrastructure, deployment pipelines, and observability that the rest of engineering relies on every day.
This is an individual contributor role focused on:
- Owning infrastructure, CI/CD, and automation initiatives end to end; from design through rollout, adoption, and iteration
- Setting the patterns and standards that keep our systems reliable, observable, and secure as we scale
- Reducing operational toil and raising the bar on developer productivity across engineering
Infrastructure & Automation
- Design, build, and maintain cloud infrastructure on Google Cloud Platform using Terraform, and help define the patterns and standards the team builds on
- Own and improve CI/CD pipelines in GitHub Actions to make deployments fast, safe, and repeatable
- Identify and eliminate sources of operational toil, building automation and tooling that scales across engineering
- Establish meaningful monitoring, dashboards, and alerting in Datadog, and drive down alert noise across the team's systems
- Lead incident response within the on-call rotation, and drive postmortems and follow-ups that improve uptime for our applications
- Define and track Service Level Objectives (SLOs) for our core infrastructure and build systems, and hold the team to them
- Partner with product engineering teams on deployment pipelines, environment issues, and build troubleshooting, and proactively remove recurring friction
- Own preview and staging environments, including reliable data sync, masking, and cleanup routines so teams can test against realistic data
- Improve developer experience through better tooling, clear runbooks, and documentation
- Write well-tested, well-reviewed infrastructure code, and lead design reviews and RFCs with a pragmatic, operations-focused perspective
- Design systems and changes to meet the team's goals, weighing tradeoffs across reliability, performance, and security
- Give thoughtful code reviews and mentor other engineers through pairing, knowledge sharing, and documentation
- Languages: TypeScript, Python, Bash
- Infrastructure: Google Cloud Platform, Terraform, Cloudflare
- CI/CD: GitHub Actions
- Databases: MySQL, PostgreSQL, Redis
- Observability & Incident Management: Datadog, Sentry, Rootly
- Integrations: Athena EMR, wearable platforms, third-party healthcare APIs
- :5+ years of professional experience in DevOps, SRE, infrastructure, or backend engineering in production environments
- Hands-on experience designing and operating infrastructure in at least one cloud provider (ideally Google Cloud Platform)
- Track record of owning and shipping automation, pipelines, or infrastructure that made a team measurably more productive or reliable
- An enthusiasm for developer productivity and making our teams as impactful as possible
- Deep experience with infrastructure-as-code (ideally Terraform) and building CI/CD pipelines (ideally GitHub Actions)
- Proficient software engineering ability, and strong command of Linux
- Strong instincts for monitoring and observability, and confident debugging across logs, traces, and metrics
- Solid experience with relational databases (MySQL, PostgreSQL) and containerized workloads
- Strong grounding in reliability, performance, and security fundamentals, with the judgment to make sound tradeoffs
- Experience in healthcare, digital health, or other regulated domains (HIPAA, PHI, SOC 2, etc.)
- Experience with containers and orchestration (Docker, Kubernetes)
- Exposure to leading incident response, on-call, and postmortem practices
- Experience with database migrations or managing multiple environments at scale
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search