Site Reliability Engineer
Indexed description
Trading Infrastructure is a global organization of Engineers who architect, build and maintain our world-class infrastructure. From colo design/implementation, to optimizing our exchange connectivity, to building world class low latency Wide Area Networks, we leverage research and automation to consistently adapt and innovate our infrastructure to scale and drive our trading and evolving business.
We are looking for exceptional talent that can work closely and collaborate effectively across multiple teams spread around the globe and continue to take our infrastructure to the next level.
What You'll Do:
- Develop deep technical expertise in your assigned product area and tech stack.
- Own production deployment, configuration, and release processes
- Drive performance, reliability, and operability through continuous improvement
- Build and maintain production tooling that supports deployment, orchestration, monitoring, and system diagnostics
- Define and maintain observability, SLI/SLOs, and performance metrics in partnership with product owners
- Leverage metrics and capacity planning to ensure scalability and uptime
- Collaborate across engineering teams to troubleshoot and resolve complex production incidents
- Lead and coordinate incident response, root cause analysis, and post-mortems
- Influence architecture and promote best practices by aligning with global SRE teams
- Document processes and procedures; provide mentorship and cross-training to peers.
- Actively manage operational risk for production changes
- Other duties as assigned or needed
- Degree in Computer Science, a related field, or equivalent professional experience
- At least 5+ years of relevant work experience in an IT ops role, such as DevOps, SRE, Linux Systems Engineering, or Network Engineering
- At least 5+ years of experience using systems programming language (C/C++/Golang/Rust)
- A rigorous, detail-oriented approach to operations
- Strong understanding of the Linux operating system, including network and system configuration, kernel internals, scheduling, performance tuning
- Strong understanding of networking concepts such as routing, multicast, LLDP, VLANs, and Ethernet
- A deep sense of ownership and desire to meet business priorities with urgency
- Ability to handle shared operational and periodic on-call duties
- Reliable and predictable availability
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search