Back to search
Infra360 Linkedin · Posted 2mo ago

Senior DevOps Engineer

India

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

AWS / Azure / GCP | Multi-Client Ownership | Technical Leadership

About The Role

We are looking for a Senior DevOps Engineer with 5–8 years of strong hands-on production experience to independently own multiple client environments and provide technical leadership to a team of engineers.

In this role, you will typically manage 2–3 client environments, drive cloud and infrastructure architecture decisions, troubleshoot complex production challenges, and lead initiatives across reliability, security, automation, performance, and cost optimization.

You will also provide technical direction and mentorship to 3–5 engineers, ensuring high-quality delivery, strong engineering practices, and effective execution.

The role requires strong architectural thinking, engineering judgment, production troubleshooting skills, and the ability to communicate effectively with both technical teams and client stakeholders.

Key Responsibilities

  • Cloud Infrastructure & Architecture
  • Design, deploy, manage, and optimize production cloud environments across AWS, Azure, and/or GCP.
  • Design and review highly available, scalable, secure, and cost-efficient cloud architectures.
  • Work extensively with core cloud services including compute, storage, databases, networking, IAM, and load balancers.
  • Own production workloads on EKS, AKS, GKE, including cluster design, upgrades, scaling, troubleshooting, and maintenance.
  • Evaluate architectural options and recommend solutions based on client requirements, technical constraints, and business objectives.
  • Identify technical debt, architectural risks, and opportunities to improve reliability, scalability, security, and operational efficiency.
  • Implement and improve backup, disaster recovery, high availability, and business continuity practices.
  • Kubernetes, Infrastructure as Code & Automation
  • Design and operate production Kubernetes environments, addressing challenges related to networking, scheduling, scaling, resource utilization, availability, and application behavior.
  • Establish reusable Kubernetes and infrastructure patterns across client environments.
  • Develop and maintain infrastructure using Terraform and Infrastructure as Code best practices.
  • Build reusable IaC modules and standards to improve consistency, scalability, and operational reliability.
  • Design and implement scalable CI/CD and GitOps workflows using tools such as ArgoCD, Flux, Spinnaker, or similar platforms.
  • Automate operational processes using Bash, Python, and other appropriate scripting or automation tools.
  • Identify and eliminate repetitive operational work through automation and engineering improvements.
  • Reliability, Monitoring & Production Operations
  • Own production reliability and operational excellence across assigned client environments.
  • Lead troubleshooting of complex infrastructure, Kubernetes, networking, and application performance issues.
  • Configure and improve monitoring, logging, alerting, and observability using tools such as Prometheus, Grafana, Coralogix, New Relic, Datadog, CloudWatch, or equivalent platforms.
  • Define appropriate SLIs, SLOs, dashboards, and alerting standards for production environments.
  • Lead production incident response, root cause analysis, and preventive remediation initiatives.
  • Identify systemic reliability risks and drive improvements to prevent recurring incidents.
  • Security & Compliance
  • Implement cloud security best practices across infrastructure and production environments.
  • Apply principles of IAM, RBAC, least privilege, secrets management, vulnerability management, and OS hardening.
  • Identify security gaps and incorporate security considerations into infrastructure and architecture decisions.
  • Work with engineering teams to improve the overall security posture of client environments.
  • Cloud Cost & Performance Optimization
  • Drive cloud cost and resource optimization initiatives across client environments.
  • Identify cloud waste through right-sizing, capacity planning, workload optimization, and appropriate cloud pricing models.
  • Balance cost, performance, reliability, security, and scalability when making technical decisions.
  • Contribute to FinOps practices and help clients achieve sustainable cloud cost efficiency.
  • Client & Stakeholder Leadership
  • Own the technical relationship for approximately 2–3 client environments.
  • Lead technical discussions with client engineering teams, architects, and leadership stakeholders.
  • Translate business requirements into practical, scalable, and maintainable technical solutions.
  • Present architecture recommendations, technical risks, trade-offs, and improvement plans to clients.
  • Handle technical escalations and production incidents with clear ownership and proactive communication.
  • Challenge inefficient or technically risky approaches and recommend better alternatives.
  • Build client trust through technical credibility, effective communication, and predictable delivery.
  • Team Leadership & Mentoring
  • Provide technical leadership to a team of 3–5 engineers across assigned client environments.
  • Plan and delegate work based on technical capability, priorities, client requirements, and business impact.
  • Review technical implementations and ensure adherence to engineering standards and best practices.
  • Mentor engineers and actively contribute to their technical development.
  • Identify capability gaps and create opportunities for knowledge sharing and skill development.
  • Provide technical guidance during complex incidents, design discussions, and implementation challenges.
  • Maintain accountability for technical quality and delivery standards.

Engineering Expectations

  • Make independent and well-reasoned technical decisions in complex or ambiguous production situations.
  • Evaluate trade-offs across reliability, scalability, security, performance, cost, and complexity.
  • Demonstrate strong ownership of production systems and follow issues through to resolution.
  • Establish and continuously improve engineering and operational standards.
  • Proactively identify technical debt, operational risks, and improvement opportunities.
  • Approach problems with a structured, analytical, and solution-oriented mindset.

Required Skills & Experience

Must Have

  • 5–8 years of hands-on experience in DevOps, Cloud Engineering, SRE, or a similar role.
  • Strong production experience with AWS, Azure, and/or GCP.
  • Strong hands-on experience with Kubernetes, preferably EKS, AKS, or GKE.
  • Strong experience with Terraform and Infrastructure as Code.
  • Strong understanding of CI/CD and GitOps practices.
  • Strong Linux and networking fundamentals.
  • Proven experience troubleshooting complex production environments and handling incidents.
  • Strong scripting and automation skills using Bash, Python, or equivalent.
  • Good understanding of cloud security fundamentals.
  • Experience designing and reviewing production architectures.
  • Demonstrated ownership of production environments and technical initiatives.
  • Experience mentoring or providing technical leadership to engineers.
  • Strong written and verbal communication skills.
  • Experience working with multiple clients, products, or production environments simultaneously.

What Success Looks Like

Within The Role, You Will Be Expected To

  • Independently own 2–3 client environments and their technical outcomes.
  • Lead and mentor 3–5 engineers while maintaining high engineering standards.
  • Resolve complex production incidents and technical escalations effectively.
  • Review and continuously improve client architectures.
  • Make sound technical decisions with minimal supervision.
  • Drive measurable improvements in reliability, security, automation, performance, and cloud cost.
  • Lead technical conversations confidently with client stakeholders.
  • Identify risks and improvement opportunities proactively rather than responding only to incidents.
  • Raise the technical maturity and engineering capabilities of the team.
Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search