Senior Site Reliability Engineer
Indexed description
We’ve helped hundreds of companies like UPS, CLEAR, Stitch Fix, GoPuff, Fetch, and sweetgreen to hire, onboard, and manage over 14 million workers in more than 75 countries.
In 2022, we closed $185M in our Series C, led by SoftBank and B Capital.
Join our growing team of highly collaborative, ambitious, and forward-thinking Fountaineers as we empower our hundreds of customers and millions of frontline workers around the world.
Let’s elevate frontline work together.
As a Site Reliability Engineer (SRE) at Fountain, you’ll own the reliability, scalability, and operational excellence of the systems that power our platform. You’ll partner closely with Engineering, Security, and Product to build resilient infrastructure, improve developer experience, and raise the bar on observability and incident response.
What You'll Be Doing
- Own and improve platform reliability (SLOs/SLIs), capacity planning, and production readiness
- Design, build, and maintain Kubernetes-based infrastructure and deployment workflows
- Operate and evolve our AWS footprint (networking, compute, storage, IAM) with a security-first mindset
- Improve our CD/GitOps practices (ArgoCD) and deployment safety (progressive delivery, rollbacks, guardrails)
- Build autoscaling strategies for services and workloads (KEDA where appropriate)
- Lead incident response: on-call, triage, mitigation, postmortems, and preventative follow-through
- Strengthen observability across services: metrics, logs, traces, and alerting (OpenTelemetry + dashboards)
- Partner with application teams to tune performance, reduce toil, and improve operational maturity
- Improve infrastructure-as-code practices and maintain Terraform modules and environments
- Contribute to evaluating and integrating AI tooling and MCP tools to accelerate operational workflows
- 5+ years of experience in site reliability engineering
- Experience owning reliability for production systems — defining SLOs or running error budgets
- Deep hands-on experience with Kubernetes/Helm on EKS in production
- Experience with Cloudflare (CDN, WAF, etc)
- Experience with CI/CD or general GitOps deployment pattern experience
- Has led or significantly contributed to incident response and postmortem processes
- Working knowledge of AWS core services (networking, IAM, compute) beyond just EKS
- Experience evaluating or building AI-assisted ops tooling (MCP, agentic runbooks)
- OpenTelemetry / observability pipeline design
- ArgoCD, KEDA autoscaling
- Experience with Terraform
Fountain offers an incredibly unique work environment. We employ a diverse team all over the world. Each Fountaineer is given the freedom to do their best work from wherever they choose. We also understand the importance of in-person connections and hold in-person meetings with your team and meet annually as an organization to build our relationships and focus on the future of moving Fountain Forward.
The benefits we offer in the United States include competitive health plans and a retirement plan. Some Fountain-wide perks offered to all employees across the globe include a flexible vacation policy, paid holidays, monthly lunch stipends, annual allowances for ongoing education related to your profession and career advancement, along with home office, cell phone, and wellness reimbursements. Fountain is a global employer, so some benefit offerings will vary from country to country.
Fountain is proud to be an equal opportunity workplace. We welcome applicants of any educational background, gender identity and expression, sexual orientation, religion, ethnicity, age, socioeconomic status, disability, and veteran status.
By submitting an application, you confirm that you have read our Privacy Policy and agree that we may process and retain your personal data for the purpose of recruitment in accordance with applicable data protection laws.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search