Application Site Reliability Engineer (SRE)
Indexed description
This is a hands-on role for an engineer who enjoys solving production challenges, improving observability, automating operations, and building resilient systems where uptime directly impacts customer experience.
Position Details
TeamPlatform & Production Reliability
LocationRemote (Americas, LatAm preferred)
Working HoursAmericas time zones (UTC-3 to UTC-8)
On-callRotation aligned with the London trading day
Employment TypeFull-time, Permanent
Experience LevelMid-Level (3-5 years)
Technology Stack.NET/C#, Windows Server, AWS, Aurora PostgreSQL, Prometheus, Grafana, Terraform
About The Role
Our trading platform powers every customer interaction, making reliability a first-class product concern. You will be responsible for maintaining and improving the operational reliability of our .NET/C# services on Windows, ensuring they remain highly available, observable, and resilient.
You'll collaborate closely with software engineers to improve monitoring, deployment safety, automation, fault isolation, and incident response, while driving continuous improvements in platform reliability and operational excellence.
What You'll Do
- Participate in the on-call rotation for production trading systems and lead incident response during service disruptions.
- Investigate production incidents, perform root cause analysis, and implement preventive actions to eliminate recurring issues.
- Build and maintain Grafana dashboards, Prometheus alerts, and operational health views across applications, infrastructure, and databases.
- Instrument .NET services to improve telemetry, metrics, logging, and visibility into service health and customer impact.
- Define, implement, and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
- Troubleshoot issues across:
- .NET/C# applications
- Windows Server
- Aurora PostgreSQL databases
- AWS infrastructure
- CI/CD pipelines and deployments
- Improve deployment safety, release automation, and rollback strategies.
- Partner with developers to improve application operability, resilience, and fault isolation.
- Automate operational tasks through scripting and infrastructure automation.
- Create and maintain runbooks, operational documentation, and incident response procedures.
- Continuously improve monitoring, alert quality, automation, and platform reliability
- SLIs & SLOs
- Error Budgets
- Incident Response
- Root Cause Analysis (RCA)
- Alert Design
- Production Operations
- Experience supporting high-availability or low-latency financial or trading systems.
- Familiarity with MetaTrader environments or financial technology platforms.
- Experience with distributed systems and microservices.
- Knowledge of OpenTelemetry or similar observability frameworks.
- Exposure to Docker, Kubernetes, or containerized environments
- Work on mission-critical trading infrastructure that directly impacts customers.
- Solve challenging reliability and scalability problems in a real-time environment.
- Build world-class observability, automation, and deployment practices.
- Collaborate with experienced engineers in a modern engineering culture.
- Influence reliability strategy and engineering best practices across the platform.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search