SRE Operation Lead
Indexed description
Role: SRE Operation Lead
Onsite at Marlborough, MA
Full Time Direct Hire Role
Job Description
Must Have Technical/Functional Skills
· Strong Python and/or Shell scripting
· Strong programming background in Java/Spring Boot and/or .NET
· REST APIs and Microservices
· SQL and database troubleshooting
· Git/source control
· CI/CD pipelines
· Application and production debugging
· Log analysis
· Monitoring/observability tools
· Linux/Unix
· Automation and self-healing implementation
· Strong RCA/problem-solving capability
Good to Have
· New Relic
· Splunk/Scalyr
· Docker/Kubernetes
· Azure or AWS
· Jenkins/Azure DevOps
· ServiceNow
· API Gateway/API Connect
· Messaging technologies such as Kafka/RabbitMQ/Service Bus
Roles & Responsibilities
Production Troubleshooting & RCA
· Troubleshoot production issues across UI, APIs, microservices, application, database and infrastructure layers.
· Analyze logs, application errors, API failures, latency and service dependencies.
· Perform code-level debugging to identify root causes rather than limiting investigation to monitoring/log analysis.
· Develop permanent engineering fixes for recurring production problems.
· Reduce MTTR through automation and faster root-cause identification.
· Client’s current SRE direction specifically emphasizes improved RCA, impact analysis, anomaly detection and faster incident resolution
Automation & Self-Healing
· Rapidly develop Python/Shell/PowerShell or application-level automation scripts for repetitive SRE activities.
· Build self-healing and auto-remediation solutions for known production failure scenarios.
· Automate recurring operational runbooks and SOPs.
· Develop automated health checks and post-deployment validation.
· Identify repetitive manual activities and convert them into zero/minimal-touch automation.
· Integrate monitoring alerts with automated remediation where applicable.
· This directly aligns with client’s documented SRE focus on self-healing systems, auto-remediation scripts, workflow automation, automated platform health checks and automation of recurring SOPs
Microservices & Full-Stack Engineering
· Strong ability to understand microservices architecture and service-to-service interactions.
· Debug REST APIs and distributed application flows.
· Ability to navigate unfamiliar codebases and quickly understand application behavior.
· Troubleshoot issues across frontend, backend, middleware and data layers.
· Understand containerized applications and distributed systems.
· Make application/code/configuration fixes when required instead of depending entirely on development teams.
Thanks & Regards,
Vaishnav Gupta| Senior Technical Recruiter
KPG99, INC | www.kpgtech.com | MBE Certified Firm
3240 E State, St Ext | Hamilton, NJ 08619
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search