Back to search
Trading Software Start Up Linkedin · Posted 13d ago

Site Reliability Engineering Manager (Chicago)

Chicago

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Lead Site Reliability Engineer


Chicago, IL or New York, NY | Hybrid

$150,000–$250,000 Base + Discretionary Bonus


Build the Infrastructure Behind a Next-Generation Trading Platform

We are partnering with a rapidly growing financial technology company building critical infrastructure and execution technology for today's electronic markets.


This is an opportunity to join at a pivotal stage of growth and play a central role in scaling the technology organization. You will work alongside a highly entrepreneurial team and have the opportunity to build, shape, and lead the Site Reliability Engineering and Systems Administration function from the ground up.


This is not a role where you inherit a mature infrastructure team and simply maintain what already exists.


You will be trusted to define how production systems are operated, automated, monitored, secured, and scaled. You'll bring structure to a growing environment, eliminate manual and repetitive work, reduce key-person dependency, and build the engineering standards and operational discipline required to support a high-performance, mission-critical platform.

If you're an experienced SRE, Production Engineer, Systems Engineer, or Infrastructure Engineer who wants real ownership, technical influence, and the opportunity to build something, this is a rare opportunity.


Reporting directly to the CTO, this role combines hands-on engineering, technical leadership, and the chance to have a meaningful impact on the future of the platform.


The Environment

You'll support an automated, real-time trading environment with single-digit millisecond latency requirements.


The core production environment runs on bare-metal Linux infrastructure, including colocation environments where performance, availability, and operational discipline are critical. Azure is leveraged for development environments, storage, and offline research—not production trade execution.


The technology environment includes:

  • Bare-metal Linux
  • Low-latency, real-time systems
  • Colocation infrastructure
  • C++, Python, and SQL
  • PostgreSQL
  • Azure for development, storage, and research workloads
  • Infrastructure as Code and automation
  • Legacy and next-generation systems running in parallel
  • Active technology modernization, including major Linux and platform upgrades
  • Continuous rollout of new technologies and infrastructure improvements


This is an environment for engineers who enjoy understanding how systems actually work—from the application and operating system down to the network and physical infrastructure.


What You'll Do

Own Production Infrastructure

  • Own how critical production systems are deployed, operated, monitored, and recovered across colocation and cloud environments.
  • Take responsibility for infrastructure capacity, availability, resiliency, failover, and performance.
  • Define and publish the engineering standards and strategy for operating production systems.
  • Build repeatable, scalable operational processes that allow the platform and team to grow.
  • Lead major infrastructure upgrades, modernization initiatives, and technology rollouts.


Build and Lead the SRE / Systems Function

  • Build and shape the systems administration and Site Reliability Engineering function.
  • Establish how the team operates as the organization transitions from individual knowledge and manual processes toward an engineering-driven, automated, standards-based discipline.
  • Mentor and develop existing technical staff and help grow the team over time.
  • Reduce key-person risk by improving documentation, automation, runbooks, and operational processes.
  • Partner directly with technology leadership to influence infrastructure strategy and long-term technical direction.


Automation & Operational Excellence

  • Automate routine infrastructure and operational work using scripting, Infrastructure as Code, and internally developed tooling.
  • Identify recurring production issues and engineer permanent solutions rather than repeatedly firefighting the same problems.
  • Improve deployment, configuration, provisioning, and recovery processes.
  • Evaluate build-vs-buy decisions and manage relationships with infrastructure and technology vendors.
  • Drive an automation-first mindset throughout production operations.


Production Support, Incident Response & Resilience

  • Own and improve incident response, on-call processes, escalation procedures, and post-incident reviews.
  • Be accountable for production stability and continuously analyze what breaks, why it breaks, and how similar failures can be prevented.
  • Develop and maintain recovery runbooks and lead recovery and resiliency exercises.
  • Partner with development teams to improve production readiness, deployment processes, performance tracking, and capacity planning.
  • Help reduce reliance on software engineering for first-line production support.
  • Build automation and self-service capabilities that reduce manual support work for both internal teams and clients.

Infrastructure Security

  • Own key aspects of infrastructure security, including system hardening, access controls, backups, recoverability, and security incident response.
  • Ensure critical production infrastructure is operated with strong security and resiliency standards.


What We're Looking For


Required Experience

  • Strong hands-on experience administering Linux in production environments.
  • Strong understanding of networking and the ability to troubleshoot complex infrastructure and network issues.
  • Excellent scripting and automation skills.
  • Deep experience with Infrastructure as Code; we are open regarding the specific tools you've used.
  • A proven track record owning production infrastructure, production support, and incident response.
  • Experience designing and improving systems for reliability, availability, performance, and recoverability.
  • Experience identifying recurring operational problems and solving them through automation and engineering.
  • Experience managing, mentoring, or developing technical staff.
  • The ability to establish structure, standards, processes, and strategy within a growing technical function.
  • Strong communication skills and the ability to work effectively with senior stakeholders, software engineers, and a small, highly collaborative team.
  • Experience leveraging modern AI-powered engineering tools to improve productivity and automation. Experience with tools such as Claude Code is a plus.


Highly Relevant Backgrounds

We would be particularly interested in candidates with experience supporting:

  • Electronic trading platforms
  • Proprietary trading firms
  • Market makers
  • Hedge funds
  • Exchanges
  • Brokerages or financial market infrastructure
  • Fintech platforms
  • Low-latency or high-performance systems
  • Real-time, mission-critical production environments


Experience in an environment where downtime, latency, and system performance have a direct impact on the business will translate particularly well.


Nice-to-Have Experience

Experience with any of the following is a plus:

  • FIX Protocol
  • PostgreSQL
  • High-performance middleware or messaging technologies
  • AERON
  • Observability and monitoring platforms
  • Microsoft Azure
  • Azure Virtual Desktop (AVD)
  • Vendor and service-provider management
  • C++, Python, and/or SQL production environments
  • Colocation or low-latency infrastructure


Why This Opportunity?

This is a chance to join a growing financial technology company at a stage where one great infrastructure leader can make an enormous difference.

You won't be buried in a massive organization or limited to maintaining a narrow piece of the stack.

You will have the opportunity to:

  • Build and lead a function from the ground up.
  • Own infrastructure supporting real-time, high-performance systems.
  • Work directly with the CTO and have genuine influence over technical strategy.
  • Solve challenging problems involving Linux, networking, automation, resiliency, security, and performance.
  • Replace manual processes and tribal knowledge with scalable engineering solutions.
  • Lead major infrastructure modernization and technology initiatives.
  • Help shape the reliability and operational culture of a growing company.
  • Work in an entrepreneurial environment where your ideas can move quickly from concept to production.
  • Make a direct and visible impact on a mission-critical platform.



We want an engineer who wants to own the environment, raise the bar, build the function, and leave a lasting technical legacy.


Compensation

Base Salary: $150,000–$250,000 USD

Bonus: Discretionary bonus

Compensation will be commensurate with experience, technical expertise, and overall fit for the role.


We are an equal opportunity employer and are committed to building an inclusive and collaborative workplace.

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search