Production Support Engineer
Indexed description
Role: Production Support Engineer – Linux / Splunk / Dynatrace
Location: Jacksonville, FL (3 days Hybrid)
Exp.: 8-10 yrs
Mandatory skills: Dynatrace, RedHat Linux Administrator, ServiceNow ITSM, Splunk
Good to have skills: Critical Incident Response, Major incident management
Role summary:
We are seeking a proactive and detail oriented Platform Support Analyst to ensure the stability availability and performance of critical production applications The ideal candidate will lead incident triage and resolution activities coordinate major incident calls perform technical troubleshooting and provide timely stakeholder communications
Primary mandatory skills:
- 3 years of Incident Management Triage Leadership in a 24x7 production support environment
- Strong experience leading Major Incident P1P2 calls and coordinating cross-functional teams
- Hands on experience with Splunk for log analysis monitoring and troubleshooting
- Strong knowledge of Linux/Unix administration and production support activities
- Experience with Dynatrace AppDynamics Introscope or similar APM Monitoring tools
- Experience supporting Batch Processing and Job Scheduling environments
- Strong troubleshooting problem solving and root cause analysis skills
- Excellent written and verbal communication skills with experience preparing executive summaries and incident communications
- Experience with ITSM tools such as ServiceNow Remedy or similar ticketing platforms
Key responsibilities:
- Lead incident triage and resolution calls as the Incident Call Leader ensuring rapid restoration of services
- Coordinate with application infrastructure database and support teams during critical incidents
- Perform technical troubleshooting using Splunk Dynatrace and other monitoring tools
- Create update and manage incident and problem tickets throughout the incident lifecycle
- Provide timely and accurate communications to technical teams business stakeholders and leadership
- Document detailed incident timelines impact assessments and resolution notes
- Prepare executive summaries incident reports and postmortem reviews
- Ensure data accuracy and quality within incident management systems
- Identify recurring production issues and drive continuous service improvement initiatives
- Develop and maintain operational procedures runbooks and support documentation
- Monitor batch processing activities and coordinate issue resolution when failures occur
- Support 24x7 operations and participate in rotational on-call support as required
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search