Systems Integration Specialist Advisor
Indexed description
Whenever possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored to each client’s needs. While many positions offer remote or hybrid work options, these arrangements are subject to change based on client requirements. For employees near an NTT DATA office or client site, in-office attendance may be required for meetings or events, depending on business needs. At NTT DATA, we are committed to staying flexible and meeting the evolving needs of both our clients and employees. NTT DATA recruiters will never ask for payment or banking information and will only use @nttdata.com and @talent.nttdataservices.com email addresses. If you are requested to provide payment or disclose banking information, please submit a contact us.
Role Summary
We are looking for a strong, hands-on Senior Site Reliability Engineer to support and improve cloud operations for microservice-based platforms. This role requires a senior engineer who can independently manage production reliability, incident response, cloud infrastructure, automation, observability, Kubernetes operations, and CI/CD workflows across AWS and Azure environments.
The ideal candidate should be technically strong, proactive, comfortable in production support, and able to reduce operational toil through automation while improving service availability, performance, scalability, and resilience.
Key Responsibilities
- Own and improve reliability of cloud-based services and supporting infrastructure.
- Participate in on-call rotation and support production systems outside normal business hours.
- Lead incident response activities including triage, escalation, mitigation, and service restoration.
- Drive blameless postmortems and ensure corrective actions are tracked to closure.
- Design, implement, and maintain Infrastructure as Code using Terraform and tools such as Atlantis.
- Manage and enhance GitOps and deployment workflows using ArgoCD and related CI/CD tools.
- Support and improve cloud/container platforms across AWS and Azure.
- Manage Kubernetes-based workloads, containers, virtual servers, and distributed systems.
- Build automation to reduce manual effort and improve operational efficiency.
- Configure and improve monitoring, alerting, logging, diagnostics, and observability.
- Analyze performance and capacity trends to identify bottlenecks and improve scalability.
- Troubleshoot complex infrastructure, networking, application runtime, and cloud platform issues.
- Support disaster recovery planning, validation, and recovery readiness.
- Create and maintain operational runbooks, support procedures, and engineering documentation.
- Coach and guide other engineers on SRE best practices, reliability, automation, and operational excellence.
- 6–8+ years of experience as an SRE, DevOps Engineer, Infrastructure Engineer, Cloud Engineer, or Platform Engineer.
- Strong hands-on experience with AWS and Azure cloud platforms.
- Strong experience with Terraform for Infrastructure as Code.
- Experience with Atlantis, ArgoCD, or similar infrastructure/deployment automation tools.
- Strong hands-on experience with Docker and Kubernetes.
- Experience designing, maintaining, and troubleshooting complex CI/CD pipelines.
- Strong production support experience including incident management, RCA, postmortems, and runbook creation.
- Strong observability experience: monitoring, alerting, logging, diagnostics, and performance analysis.
- Good understanding of cloud networking, security, access controls, and InfoSec practices.
- Experience with version control, branching, merging, pull requests, and conflict resolution.
- Understanding of cloud cost optimization and resource utilization.
- Strong communication skills and ability to work with DevOps, Engineering, Product, and Delivery teams.
- Experience with microservice-based platforms.
- Experience with tools such as Datadog, CloudWatch, Grafana, Prometheus, Splunk, AppDynamics, or similar.
- Scripting/programming experience using Python, Bash, Go, or Java.
- Experience with SLI/SLO/SLA, error budgets, capacity planning, and resilience engineering.
- Experience with disaster recovery testing and production readiness reviews.
- Experience supporting customer-facing, high-availability platforms.
- Prior experience mentoring junior engineers or leading technical troubleshooting.
We are currently seeking a Systems Integration Specialist Advisor to join our team in Guadalajara, Jalisco (MX-JAL), Mexico (MX).
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search