Tata Consultancy Services
Linkedin · Posted 2d ago
DevOps & Site Reliability Engineer (Digital)
Continue to application
Add your email once, then Caio opens the original posting.
Indexed description
Job DescriptionMust Have Technical/Functional Skills
- Technology and Programming (Expert Level)
- Strong proficiency in Java full stack developer
- Object-Oriented programming principles and concepts
- Hands-on experience on Observability platform Dynatrace
- Hands-on experience with Spring Framework (Spring Boot, Spring MVC, Spring Security)
- Knowledge if RESTful API development
- Experience with database like Oracle, DB2, MySQL
- Proficiency in Payment Switch BASE24 EPS, C++, AS400 and Python is also added advantage
- Domain, Cloud & Platform Engineering
- Must have domain experience on Retail Point of Sale/Payment Systems/Merchandising/Inventory/Logistics area
- Expertise in Microsoft Azure, including:
- Compute (VMs, App Services, Azure Container Apps)
- Containers & Orchestration (AKS, Docker)
- Storage, Azure Key Vault, Azure Monitor, Log Analytics
- Proven experience designing enterprise grade, highly available cloud platforms
- Advanced experience with Azure DevOps and CI/CD pipeline architecture
- Strong scripting skills (PowerShell, Bash)
- GitOps concepts, branching strategies, release orchestration
- Ownership of platform reliability, resiliency, and performance
- SLIs, SLOs, SLAs
- Error budgets and reliability metrics
- Metrics, logs, traces, alerts, dashboards using Dynatrace
- Incident response leadership, RCA facilitation, and long term remediation planning
- Experience operating 99.9%–99.99% availability systems
- Secure cloud design using Key Vault, managed identities, RBAC
- Cost optimization (FinOps mindset) across cloud infrastructure
- Act as SRE Technical Architect (Should be interested to work on Implementations) for client's Retail platforms, owning reliability and stability outcomes
- Define and enforce SRE standards, best practices, and operating models
- Architect and govern highly available, scalable cloud platforms
- Lead the design and implementation of CI/CD and IaC strategies
- Establish proactive monitoring, alerting, and incident prevention mechanisms
- Own major incident leadership, RCA execution, and corrective action tracking
- Partner with application, security, and architecture teams to build reliability by design
- Drive automation to reduce toil and improve operational efficiency
- Mentor and coach SRE and DevOps engineers across teams
- Influence roadmap decisions with a reliability, scalability, and cost lens
- Discretionary Annual Incentive.
- Comprehensive Medical Coverage: Medical & Health, Dental & Vision, Disability Planning & Insurance, Pet Insurance Plans.
- Family Support: Maternal & Parental Leaves.
- In surance Options: Aut& Home Insurance, Identity Theft Protection.
- Convenience & Professional Growth: Commuter Benefits & Certification & Training Reimbursement.
- Time Off: Vacation, Time Off, Sick Leave & Holidays.
- Legal & Financial Assistance: Legal Assistance, 401K Plan, Performance Bonus, College Fund, Student Loan Refinancing.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search
Want help applying to roles like this?
Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search