MLOps Engineer
Indexed description
Working closely with ML engineers, data scientists, and platform engineers, this role will improve release velocity without compromising reliability, security, or model quality. The team supports workloads across batch inference, real-time APIs, and LLM-enabled applications.
Key Responsibilities
- Build and maintain end-to-end ML pipelines using Kubeflow, Airflow, or equivalent orchestration frameworks for training, validation, and deployment
- Automate model packaging and release workflows with Docker, Kubernetes, Helm, and GitHub Actions or GitLab CI
- Operate model serving infrastructure on AWS, GCP, or Azure using platforms such as SageMaker, Vertex AI, KServe, or NVIDIA Triton
- Implement model and data observability for latency, throughput, drift, feature quality, cost, and production performance using tools such as Prometheus, Grafana, and MLflow
- Manage experiment tracking, model versioning, artifact storage, and promotion workflows across development, staging, and production environments
- Improve platform reliability through infrastructure as code, automated testing, incident response, capacity planning, and documented runbooks
- Partner with ML and backend engineers to optimize inference performance, GPU utilization, deployment safety, and rollback procedures
- 3–8 years of experience in MLOps, ML platform engineering, DevOps, or a closely related software engineering discipline, including experience supporting production ML workloads
- Strong Python and Linux skills, with the ability to write maintainable automation, services, and deployment tooling
- Hands-on experience with Kubernetes, Docker, CI/CD, and infrastructure as code using Terraform, Pulumi, or equivalent tools
- Production experience with at least one major cloud platform: AWS, GCP, or Azure, including networking, IAM, storage, and compute services
- Working knowledge of ML lifecycle systems such as MLflow, Kubeflow, Airflow, SageMaker, Vertex AI, or comparable platforms
- Understanding of model serving patterns, feature and training data pipelines, monitoring, security, and reliability engineering principles
- Bonus: experience with GPU scheduling, distributed training, LLM inference, Ray, KServe, NVIDIA Triton, service meshes, or a degree in computer science, engineering, mathematics, or a related technical field
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search