Production Support Engineer
Indexed description
Job Title: Production Support Engineer
Location: Lafayette, Louisiana, United States
Duration: Long Term Contract
About: VLink, founded in 2006, is a leading global provider of software engineering services with next-gen technologies and best-in-class talent. Our Headquarters are in the U.S, and we have offices in 7+ countries from North America-Europe to APAC, with expansion plans in the Middle East. With over 1,000 employees working globally, VLink has helped SMBs, and large enterprises achieve their business goals, and gained the trust of Fortune-250 companies. VLink is 'Great Place to Work® Certified™' and has been a consistent winner as- Best Places to Work in CT. Trust, collaboration, and accountability are the three elements that are at the core of VLink's work culture. We value our professionals, providing comprehensive benefits and the opportunity for growth.
Position Description:
Client is seeking a Production Support Engineer who will provide production support for enterprise applications and cloud-based platforms, ensuring high availability, reliability, performance, and stability.
The role requires strong troubleshooting and coding skills across React, Java/Spring Boot, Python, APIs, microservices, databases, caching technologies, and AWS. The engineer will work closely with development, infrastructure, DevOps, and business teams to resolve production issues and continuously improve application support and operational efficiency.
This is a Full-Time, On-Site employment opportunity located in Lafayette, LA or Bloomfield, CT in a Hybrid Model.
Your future duties and responsibilities:
- Provide L2/L3 production support for enterprise applications developed using React, Python, APIs, microservices, MongoDB, Redis, and AWS services.
- Monitor application availability, performance, and reliability, and proactively identify production risks and service degradation.
- Having Coding knowledge on React UI, Java/Spring Boot backend service, integration, database, cache, and application performance issues.
- Support MongoDB operations, including connectivity, data validation, query optimization, indexing, and performance troubleshooting.
- Monitor and resolve Redis Cache issues related to availability, memory utilization, key expiration, data synchronization, and application connectivity.
- Support AWS hosted applications and deployments involving Amazon S3, CloudFront, React UI components, and related cloud services.
- Use Dynatrace and Splunk for log analysis, distributed tracing, alert investigation, performance monitoring, dashboarding, and production diagnostics.
- Manage production incidents by performing impact assessment, participating in incident bridges, coordinating resolution, completing root cause analysis, and implementing preventive actions.
- Develop Python based automation for application health checks, log analysis, alert enrichment, operational reporting, and repetitive support activities.
- Apply Generative AI, prompt engineering, RAG, and AI assisted tools to accelerate incident analysis, generate incident summaries, support troubleshooting, and improve knowledge management.
- Support application releases by completing readiness checks, validating deployments, monitoring post release performance, and coordinating rollback or remediation activities when required.
- Communicate incident status, risks, technical findings, and recovery progress clearly to business stakeholders, development teams, infrastructure teams, and leadership.
- Participate in rotational on call support and continuously improve application stability, support efficiency, monitoring coverage, and incident prevention.
Required qualifications to be successful in this role:
- At least 6–8 years of overall IT experience, including at least 4 years of L2 production support for enterprise and business critical applications.
- Having coding knowledge for supporting applications developed using React, Python, REST APIs, microservices, MongoDB, Redis Cache, and AWS services.
- Experience troubleshooting end to end production issues across the UI, backend services, APIs, integrations, databases, caching layers, and cloud infrastructure.
- Experience designing and building an enterprise alerting framework for proactive monitoring, alert correlation, notification, escalation, and incident prevention.
- Experience supporting Doctor Tools and clinician facing applications, including application availability, integrations, workflow issues, and production performance.
- Strong experience managing incident and change queues, including ticket prioritization, assignment, SLA tracking, technical analysis, change validation, stakeholder communication, and timely closure.
- Experience supporting MongoDB, including connectivity, query performance, indexing, data validation, and production troubleshooting.
- Working knowledge of Redis Cache, including availability, memory utilization, key expiration, synchronization, connectivity, and performance issues.
- Practical experience using AI tools, Generative AI, prompt engineering, and RAG for incident analysis, troubleshooting, operational automation, incident summarization, and knowledge management.
- Experience supporting, migrating, and operating AI enabled IVR and contact center solutions, including Kore.ai, Sierra AI for Health Services, Amazon Connect Outbound, Doctor Tools, and related AI tools, covering integrations, monitoring, incident resolution, production stability, and continuous improvement.
- Experience supporting AWS hosted applications involving Amazon S3, CloudFront, application deployments, monitoring, and related AWS services.
- Strong experience using Dynatrace and Splunk for log analysis, distributed tracing, dashboarding, alert investigation, performance monitoring, and root cause analysis.
- Proven experience managing major incidents, incident bridges, problem management, root cause analysis, corrective actions, and preventive measures.
- Experience developing Python automation scripts for health checks, log analysis, alert enrichment, operational reporting, and repetitive support activities.
- Experience supporting production releases, deployment validation, post release monitoring, rollback coordination, runbooks, SOPs, and knowledge documentation.
- Ability to participate in rotational on call support and communicate effectively with business stakeholders, client teams, development teams, infrastructure teams, and leadership.
Employment Practices:
VLink is an equal opportunity employer committed to fostering an inclusive environment where diversity is celebrated. All qualified applicants will be considered for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or veteran status. Employment is contingent upon successful completion of a background check.
This job application process may use AI-powered tools to assist in screening and evaluating applications based on objective, job-related qualifications. AI is used solely to support the recruitment process and all final hiring decisions are made exclusively by our human recruitment team.
Applicant information will be handled in accordance with VLink's privacy policy.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search