Site Reliability Engineer
Indexed description
Prodapt is the largest specialized player in the Connectedness industry. As an AI-first strategic technology partner, Prodapt provides consulting, business reengineering, and managed services for the largest telecom and tech enterprises building networks and digital experiences of tomorrow. A ServiceNow-invested company, Prodapt has been recognized by Gartner as a Large, Telecom-Native, Regional IT Service Provider. A “Great Place To Work® Certified™” company, Prodapt employs over 5,000 technology and domain experts across the Americas, Europe, India, Africa, & Japan. Prodapt is part of the 130-year-old business conglomerate The Jhaver Group, which employs over 32,000 people across 80+ locations globally.
We are seeking a highly technical and analytical Site Reliability Engineer (SRE) specializing in Environment Management and DevOps to join our engineering team. In this role, you will bridge the gap between development and operations by owning the health, uptime, scalability, and automated provisioning of our development, testing, staging, and production environments. You will focus on building automated guardrails, optimizing software delivery pipelines, and ensuring environmental consistency across our entire software ecosystem.
Key Responsibilities
- Environment Provisioning & Management: Standardize, build, and maintain multi-tier staging and production environments using Infrastructure as Code (IaC) templates to eliminate environment drift.
- CI/CD Pipeline Engineering: Design, implement, and optimize robust build and deployment automation pipelines to deliver reliable software updates with zero downtime.
- Observability & Monitoring: Instrument deep monitoring, logging, and tracing frameworks across all systems to establish early warning infrastructure alerts and clear system metrics (SLIs/SLOs).
- Automation & Orchestration: Eradicate repetitive manual operational tasks ("toil") by developing automation scripts and managing container orchestration platforms.
- Incident Response & Post-Mortems: Act as a critical responder for infrastructure issues, driving rapid incident resolution followed by rigorous blameless post-mortems to prevent recurrence.
- Capacity & Performance Management: Conduct environment stress testing, load balancing configuration, and resource capacity planning to ensure optimal system scaling and cost efficiency.
Required Qualifications
- Experience: Minimum of 3 to 5 years of experience operating in an SRE, DevOps, or Environment Management capacity.
- Education: Bachelor’s degree in computer science or a related technical field.
- Infrastructure as Code (IaC): Advanced proficiency with automation tools like Terraform, OpenTofu, Ansible, CloudFormation, or Bicep.
- CI/CD Expertise: Proven hands-on experience structuring delivery pipelines using Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
- Containerization: Strong working knowledge of Docker and container orchestration via Kubernetes or managed Kubernetes ecosystems.
- Scripting Languages: Strong proficiency in writing clean automation code using Python, Bash, or Go.
- Systems & Cloud Knowledge: Deep understanding of Linux/Windows OS internals, networking fundamentals (DNS, load balancers, HTTP/S, VPCs), and cloud platforms (AWS, Azure, or GCP).
- Work Arrangement: Must be fully available to work physically on-site.
Preferred Qualifications
- Certifications such as Certified Kubernetes Administrator (CKA) or professional-level Cloud Solutions Architect/DevOps certifications.
- Experience deploying enterprise observability stacks such as Prometheus, Grafana, ELK (Elasticsearch/Logstash/Kibana), or Datadog.
- Solid understanding of GitOps methodologies using tools like ArgoCD or Flux.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search