Bajaj Broking
Linkedin · Posted 8d ago
Site Reliability Engineer
Continue to application
Add your email once, then Caio opens the original posting.
Indexed description
About The RoleAt Bajaj Broking, high availability isn't just an internal KPI—it’s the backbone of every trade, order execution, and portfolio update across millions of retail investors. During high-volatility market opens, our systems process massive spikes in concurrent transactions. Zero downtime is our baseline.We are looking for a Lead Site Reliability Engineer (SRE) with 8–10 years of engineering experience to own the resilience, scalability, and performance of our core trading platforms. You will bridge the gap between software engineering and systems operations—building self-healing infrastructure, automating away operational toil, and maintaining ultra-low latency infrastructure.
What You’ll Do
- Architect for Scale: Design, deploy, and manage self-healing, high-concurrency cloud infrastructure and multi-tenant Kubernetes clusters.
- Eliminate Toil: Drive an "Automation-First" culture using Terraform, Ansible, and Python/Go to automate infrastructure provisioning, auto-scaling, and failovers.
- Master Observability: Build full-stack observability pipelines (Prometheus, Grafana, Datadog, ELK) to capture high-cardinality metrics, tracing, and log aggregation before issues hit production.
- Define Reliability Standards: Establish and enforce SLIs, SLOs, and Error Budgets across microservices teams to strike the right balance between rapid deployment and platform stability.
- Incident Commander & RCA: Lead high-severity incident responses, conduct blameless Root Cause Analyses (RCAs), and build long-term engineering fixes to eliminate recurring failure modes.
- CI/CD Optimization: Optimize continuous deployment pipelines (GitHub Actions, GitLab CI, ArgoCD) for zero-downtime, continuous release cycles.
- Disaster Recovery (DR) & Chaos Engineering: Conduct chaos testing and design multi-region disaster recovery protocols to guarantee business continuity under any scenario.
- 8–10 years of hands-on experience in SRE, Platform Engineering, or Cloud Infrastructure handling mission-critical, large-scale systems.
- Containerization & Orchestration: Deep expertise with Kubernetes (EKS/GKE/Self-hosted) and Docker in production.
- Infrastructure as Code (IaC): Advanced hands-on mastery of Terraform and modular automation tools.
- Cloud Mastery: Strong expertise across AWS, Azure, or GCP core services, networking (VPCs, BGP, DNS, Load Balancers), and IAM security models.
- Coding & Scripting: Strong programming ability in Python, Go, or Bash to build custom operators, internal tools, and automation scripts.
- Observability Expert: Practical experience setting up distributed tracing, APM, and automated alerting frameworks.
- Fintech Experience: Prior exposure to high-frequency trading platforms, broking systems, payment gateways, or banking backends.
- Service Mesh: Exposure to Istio or Linkerd.
- Certifications: CKA/CKAD, AWS Solutions Architect Professional, or GCP Cloud Engineer.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search
Want help applying to roles like this?
Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search