Site Reliability Engineer
Indexed description
Site Reliability Engineer
Location: Chicago, IL (Hybrid)
A leading proprietary trading firm is seeking a Site Reliability Engineer to support and enhance the reliability, performance, and scalability of critical trading infrastructure.
Responsibilities
- Build and maintain reliable, high-performance infrastructure supporting trading and research systems.
- Develop automation and operational tooling using Python.
- Monitor system health, troubleshoot production issues, and drive performance improvements.
- Support and optimize HPC clusters and distributed compute environments.
- Improve observability, incident response, and operational processes.
- Collaborate with software engineers to enhance platform reliability and efficiency.
Requirements
- 4+ years of experience in Site Reliability Engineering, Production Engineering, or Infrastructure Engineering.
- Strong Python programming and automation skills.
- Experience supporting High Performance Computing (HPC) environments.
- Strong Linux administration and troubleshooting experience.
- Excellent problem-solving skills in mission-critical environments.
Preferred
- Networking knowledge (TCP/IP, routing, switching, network troubleshooting).
- Experience supporting low-latency or high-throughput systems.
- Background in proprietary trading, quantitative finance, or other performance-sensitive environments.
This is an opportunity to work on mission-critical systems at the heart of a high-performing trading organization, alongside talented engineers in a fast-paced and highly technical environment.
This is a hybrid role based out of the firms Chicago office requiring 3 days of onsite work per week, and a rotational on call schedule.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search