Back to search
Zorba AI Linkedin · Posted 7d ago

Site Reliability Engineer (SRE)

India

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Job Summary

We are seeking an experienced Site Reliability Engineer (SRE) to join our engineering team. The ideal candidate will have strong expertise in Azure/GCP cloud platforms, DevOps practices, Kubernetes, Infrastructure as Code, CI/CD, monitoring, and production support. The role focuses on improving application reliability, automating operations, optimizing system performance, and ensuring high availability for enterprise applications in a microservices environment.

The candidate should possess strong troubleshooting skills, experience with modern cloud-native architectures, and a passion for automation and continuous improvement.

Key Responsibilities Site Reliability Engineering

  • Implement Site Reliability Engineering (SRE) best practices to improve application availability, scalability, and reliability.
  • Monitor production systems and ensure service health through proactive monitoring and alerting.
  • Define and maintain SLI, SLO, and Error Budgets.
  • Participate in production incident management, root cause analysis (RCA), and post-incident reviews.
  • Provide L2/L3 production support and participate in on-call rotations.

Cloud Infrastructure

  • Deploy, manage, and maintain cloud infrastructure on Microsoft Azure and/or Google Cloud Platform (GCP).
  • Manage Kubernetes environments such as AKS and GKE.
  • Implement Infrastructure as Code (IaC) using Terraform and Ansible.
  • Automate cloud provisioning, deployments, and infrastructure management.

DevOps & CI/CD

  • Design and maintain CI/CD pipelines using GitHub Actions.
  • Implement automated build, testing, deployment, and release processes.
  • Improve deployment reliability through automation and DevOps best practices.
  • Collaborate with development teams to enhance release quality and deployment efficiency.

Monitoring & Reliability

  • Configure and maintain monitoring, logging, and alerting solutions.
  • Analyze application and infrastructure logs using Splunk, Dynatrace, Grafana, or similar tools.
  • Develop dashboards and monitoring metrics to ensure system reliability.
  • Improve monitoring through automation and preventive measures.

Automation

  • Develop automation scripts using Python or C#.
  • Automate operational tasks, deployments, housekeeping, and infrastructure management.
  • Improve operational efficiency by reducing manual interventions.

Production Support

  • Troubleshoot complex production issues across distributed systems.
  • Perform code analysis, log analysis, and performance tuning.
  • Collaborate with engineering teams to resolve critical production incidents.
  • Maintain production stability while minimizing downtime.

Collaboration

  • Work closely with cross-functional teams including Development, DevOps, QA, and Product teams.
  • Participate in Agile ceremonies and SRE governance activities.
  • Share operational best practices and contribute to continuous improvement initiatives.

Mandatory Skills

  • 5+ years of experience in Site Reliability Engineering, DevOps, Production Support, or Platform Engineering.
  • Strong experience with Microsoft Azure and/or Google Cloud Platform (GCP).
  • Hands-on experience with Kubernetes (AKS/GKE).
  • Strong knowledge of DevOps practices and CI/CD pipelines.
  • Experience building CI/CD workflows using GitHub Actions.
  • Infrastructure as Code using Terraform and/or Ansible.
  • Experience with Splunk, Dynatrace, Grafana, or similar monitoring tools.
  • Strong experience with microservices architecture.
  • Experience with ServiceNow or other ITSM tools.
  • Knowledge of ITIL processes.
  • Strong troubleshooting and production support experience.
  • Experience working with SLI, SLO, Error Budget, and reliability engineering practices.
  • Strong understanding of API-based architectures.
  • Experience with desktop and mobile application support.

Skills: sre,devops,azure,reliability engineering
Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search