Back to search
Wall Street Consulting Services LLC Linkedin · Posted 2d ago

Level 2 Production Services Application Support

Lake Mary, Florida, United States

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Title: Level 2 Prod Services Application Support

Location: New York (NY), Lake Mary (FL)

Hybrid work ( 3 days from Office is Mandatory and 2 days remote)

Key Responsibilities

  1. Monitor, troubleshoot, and resolve Level 2 production incidents across AI platforms, cloud infrastructure, data pipelines, model-serving environments, and associated services.
  2. Provide hands-on support for deployment, orchestration, and operational management of AI/ML workloads across cloud-native environments.
  3. Build and maintain monitoring, alerting, and observability capabilities for infrastructure, applications, data pipelines, model operations, and distributed compute workloads.
  4. Perform root cause analysis for production issues, implement permanent fixes, and drive problem-management activities to improve service reliability.
  5. Collaborate with DevOps, MLOps, data engineering, platform engineering, and application teams to maintain and enhance CI/CD and deployment automation.
  6. Develop and support automation, self-service tooling, recovery mechanisms, and self-healing controls to reduce manual operational effort.
  7. Monitor and troubleshoot batch processes, workflow orchestration, data ingestion, model training, and model deployment failures.
  8. Apply ITIL-based incident, problem, change, and release management processes to support stable production operations.
  9. Analyse support-ticket and incident trends; recommend and implement operational improvements, including AI-driven automation where appropriate.


Qualifications & Skills

Mandatory:

  1. 3–5 years of experience in Level 2 application, platform, DevOps, or production support roles.
  2. Strong hands-on experience with UNIX/Linux, SQL, and shell or Python scripting.
  3. Experience troubleshooting cloud-native applications, distributed systems, containers, and Kubernetes-based environments.
  4. Working knowledge of CI/CD pipelines, deployment automation, and source-control platforms such as GitLab.
  5. Experience with monitoring, logging, and observability tools such as Splunk, Grafana, AppDynamics, Prometheus, or similar tools.
  6. Understanding of incident management, root cause analysis, problem management, and ITIL support processes.
  7. Strong analytical and problem-solving skills, with a client-service mindset.
  8. Ability to troubleshoot data-pipeline, workflow, API, and production deployment issues.



Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search