Senior Site Reliability Engineer (SRE / Backend) f/m/d
Indexed description
This is a hands-on role with real ownership — you'll set the standards for how we run production rather than inherit someone else's. The split is roughly 70% infrastructure and reliability work, 30% backend development. Most of your time goes to the platform, but you'll be comfortable dropping into the application code to debug a slow query, fix a Lambda, or ship an endpoint alongside the product team.
Because we handle sensitive mental health data, reliability and security aren't abstract goals here. When someone books a session in a difficult moment, our platform needs to work.
🚀 What You'll Do
- Own our AWS infrastructure end to end — Lambda, ECS Fargate, SQS, SNS, EventBridge, SES, Cognito, DynamoDB, RDS Postgres, and DMS
- Manage everything as code in Terraform, with well-designed modules, clean state management, and a solid review workflow
- Build and maintain CI/CD pipelines with safe rollout and rollback across web, mobile backends, and infrastructure
- Design our event-driven services for resilience: retries, dead-letter queues, idempotency, graceful degradation
- Own our Datadog and Sentry setup — define SLOs, build dashboards, and keep alerting actionable instead of noisy
- Lead incident response and run blameless postmortems that actually change how we build
- Harden our security posture: IAM, secrets management, network boundaries, Cognito auth flows, and vulnerability remediation
- Protect sensitive health data and support our GDPR and compliance requirements
- Monitor and optimize AWS spend without compromising reliability
- Contribute to backend development — APIs, event consumers, data pipelines, and Postgres and DynamoDB performance
- Participate in architectural discussions and mentor engineers on operational excellence
- 5+ years in SRE, DevOps, platform, or backend engineering, with real production ownership
- Deep AWS experience across serverless and containers — Lambda, ECS Fargate, and debugging both under pressure
- Strong Terraform skills, including module design and managing state across multiple environments
- Hands-on experience with event-driven architecture (SQS, SNS, EventBridge) and a healthy respect for its failure modes
- Solid PostgreSQL: query tuning, indexing, connection management, and zero-downtime migrations
- Production experience with Datadog or a comparable observability platform (Grafana, New Relic, Honeycomb)
- Comfortable writing production backend code in [Python / Node.js / Go]
- Genuine on-call and incident response experience — you've led an incident and written the postmortem
- Strong security fundamentals: IAM, least privilege, secrets, network isolation, common web vulnerabilities
- Pragmatic about complexity — you reach for the simplest thing that meets the reliability bar
- Effective communicator who can explain a technical tradeoff without jargon
- AWS DMS or other data migration and replication tooling
- Compliance experience (GDPR, SOC 2, ISO 27001)
- Experience in a B2B SaaS environment or healthcare-related product
- Real ownership: a small team, short feedback loops, and no layers of approval between you and production
- Work that matters: the reliability you build directly affects people reaching for mental health support
- Free access to the nilo app (incl. family support)
- A dedicated learning budget for your personal and professional development
- Work abroad for up to 90 days per year (within the EU)
- Hybrid working model: 2 days/week from the office, 3 days from home
- Urban Sports Club membership at a discounted price
- Equity options: you benefit from any increase in nilo's valuation that you've helped to create
- Regular team and company events
- Bring your dog to work: we have 4 office dogs
We would like to encourage you to apply even if the technical requirements cannot be met 100%.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search