Senior Site Reliability Engineer
Indexed description
As a Site Reliability Engineer at Runware, you will help ensure these systems remain reliable, performant and resilient as we scale. This is a highly technical, hands-on role working across software, infrastructure and production operations to improve observability, reduce incidents, eliminate operational toil and build lasting improvements across complex distributed systems.
What You'll Do
- Own and improve the reliability, availability and performance of critical production services across the Runware platform
- Define and evolve our reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standards
- Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads, participating in our engineering on-call rotation
- Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting engineering improvements
- Reduce operational toil through automation, automated remediation and improvements to deployment safety, recovery and system resilience
- Work closely with Engineering and DevOps teams on capacity planning, performance, scaling and architectural improvements as the platform grows
- Have strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering or similar role
- Have a strong understanding of distributed systems and are comfortable debugging across applications, databases, queues, containers, networking and infrastructure
- Have experience designing and operating observability systems using metrics, logs and distributed tracing
- Understand SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management and reducing operational toil
- Have experience with Kubernetes, containers, IaC and automated deployment practices, alongside the ability to write software and automation using languages such as Python, Go or PHP
- Take strong ownership of production problems and are comfortable participating in an engineering on-call rotation, taking issues from initial investigation through to long-term remediation
- Experience operating high-throughput or low-latency APIs and distributed systems
- Experience with bare-metal infrastructure, GPU environments or AI and ML workloads
- Experience with RabbitMQ or other distributed messaging and queueing systems
- Experience operating MySQL, Redis, ClickHouse or similar production data systems
- Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure environments
- Experience building automated scaling, capacity management or self-healing systems
Our release cycles are fast and intense, but they're followed by real downtime. After big pushes we expect the team to unplug, recharge, and come back ready & stronger than ever for the next leap.
- Generous paid time off - vacation, sick days, public holidays
- Meaningful stock options - share in the upside you create
- Remote-first setup - work from home anywhere we can employ you
- Flexible hours - own your schedule outside core collaboration blocks
- Family leave - paid maternity, paternity, and caregiver time
- Company retreats - twice-yearly gatherings in inspiring locations
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search