Back to search
Holiday Robotics Linkedin · Posted 9d ago

ML Systems / MLOps Engineer: Robotics

Seoul Incheon

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

← Back to Careers

ML Systems / MLOps Engineer: Robotics

Seoul Gangnam (On-site)

  • Full-time·Talent Pool / Year-round Interest

About The Role

Robot models improve only as fast as the systems that train, evaluate, release, and observe them. We are building a talent pool of ML systems and MLOps engineers to create the production backbone for Holiday Lab, our platform for turning real-world data into reusable robot skills.

You will make experiments reproducible, GPU resources efficient, models traceable, and deployments safe across cloud, lab, simulation, and robot environments. The role is measured by researcher velocity and production reliability—not infrastructure for its own sake.

Key Areas

Distributed trainingExperiment managementModel registryEvaluation automationGPU orchestrationEdge deploymentObservability

What You'll Do

  • Build reproducible training, evaluation, packaging, and release pipelines for perception, robot learning, VLA, and world models.
  • Operate GPU clusters and schedulers; improve utilization, throughput, queueing, checkpointing, and failure recovery.
  • Create experiment tracking, dataset/model lineage, registry, artifact, and configuration systems.
  • Develop deployment and rollback workflows for simulators, lab machines, and robot compute.
  • Implement monitoring for data drift, model regressions, latency, resource use, and task-level robot outcomes.

Required Qualifications

  • Professional experience building ML platforms, distributed systems, data infrastructure, or production ML services.
  • Strong Python and Linux skills, including debugging processes, networking, storage, and GPU workloads.
  • Experience with containers, CI/CD, workflow orchestration, experiment tracking, and model or artifact registries.
  • Experience operating cloud or on-prem compute and handling reliability, security, observability, and cost tradeoffs.
  • Ability to work closely with researchers and convert recurring friction into simple, durable platform capabilities.

Preferred Qualifications

  • Experience with Kubernetes, Slurm, Ray, Kubeflow, Airflow, Argo, MLflow, Weights & Biases, or equivalent systems.
  • Experience with multi-node GPU training, high-throughput storage, data caching, and checkpoint optimization.
  • Experience deploying models to NVIDIA edge devices, robot computers, or intermittent-connectivity environments.
  • Experience with ROS/ROS2, simulation farms, hardware-in-the-loop, or robotics release workflows.
  • Experience supporting foundation-model training and evaluation at meaningful scale.

What We Offer

  • Competitive compensation based on experience, level, and technical impact.
  • The tools, compute, robots, and equipment needed to do your best work.
  • The opportunity to build, test, and deploy directly on FRIDAY and the Holiday Robotics full stack.
  • Close collaboration with researchers and engineers across hardware, control, simulation, learning, perception, data, and operations.

This is an evergreen talent-pool posting. We review profiles continuously, and the timing and level of a formal hiring process may depend on team priorities and available openings. If your work matches our direction, we would still like to hear from you.

Apply

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search