Back to search
ASM Tech Solutions Linkedin · Posted 7d ago

L2 Production/Platform Support Engineer

Lake Mary, Florida, United States

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Job Title: AI/ML Platform L2 Support

Location: Lake Mary, FL or 240 NY Office

Work Arrangement: Hybrid - 3 Days Onsite / 2 Days Remote

Position Overview

We are looking for an experienced L2 Production/Platform Support Engineer to support AI/ML platforms, cloud infrastructure, data pipelines, model-serving environments, and related production services.

The ideal candidate will have strong hands-on experience with Linux/UNIX, SQL, Python/Shell scripting, Kubernetes, cloud-native applications, CI/CD, monitoring, and production incident management. Exposure to MLOps and AI/ML platforms is highly preferred.

Key Responsibilities

  • Monitor, troubleshoot, and resolve Level 2 production incidents across AI/ML platforms and cloud environments.
  • Support deployment, orchestration, and operational management of AI/ML workloads.
  • Build and maintain monitoring, logging, alerting, and observability solutions.
  • Perform root cause analysis (RCA) and implement permanent fixes.
  • Collaborate with DevOps, MLOps, Data Engineering, Platform Engineering, and Application teams.
  • Support CI/CD pipelines and deployment automation.
  • Troubleshoot data pipelines, workflow orchestration, model training, and model deployment issues.
  • Develop automation, self-service tools, recovery mechanisms, and self-healing solutions.
  • Follow ITIL-based incident, problem, change, and release management processes.
  • Analyze incident trends and identify opportunities for automation and operational improvement.

Mandatory Skills

  • 3-5 years of L2 Application, Platform, DevOps, or Production Support experience.
  • Strong UNIX/Linux and SQL.
  • Shell scripting or Python.
  • Kubernetes and container-based environments.
  • Cloud-native application and distributed-system troubleshooting.
  • CI/CD and GitLab.
  • Monitoring/observability tools such as Splunk, Grafana, AppDynamics, or Prometheus.
  • Incident management, RCA, problem management, and ITIL.
  • Experience troubleshooting APIs, data pipelines, workflows, and production deployments.

Good to Have

  • MLOps / AI/ML production support
  • Model deployment and monitoring
  • AWS, Azure, or GCP
  • Docker
  • Terraform / Ansible
  • Kafka / MQ
  • Control-M or similar schedulers
  • Self-healing and auto-remediation
  • Resilience, backup, and disaster recovery
  • AI-driven operational automation
Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search