MLOps Engineer
Indexed description
The team works closely with machine learning engineers and data scientists to bridge the gap between experimental notebooks and robust, distributed production pipelines.
Key Responsibilities
- Design, build, and maintain scalable MLOps infrastructure for training, fine-tuning, and serving large-scale machine learning models
- Implement robust CI/CD pipelines for ML workflows using tools like GitHub Actions, Argo Workflows, or Kubeflow
- Manage and optimize model serving endpoints using Kubernetes, Triton Inference Server, or vLLM to ensure low latency and high throughput
- Set up comprehensive monitoring and observability for deployed models to track GPU utilization, latency, and model drift
- Automate data ingestion, feature stores, and validation pipelines to ensure high data quality across training and inference environments
- Collaborate with security and platform engineering teams to ensure secure, compliant, and cost-effective cloud resource usage
- 3–6 years of experience in MLOps, DevOps, or machine learning engineering with a strong focus on infrastructure automation
- Deep expertise in Kubernetes, Docker, and container orchestration in cloud environments (AWS, GCP, or Azure)
- Hands-on experience with ML tooling such as Kubeflow, MLflow, Airflow, or Feast
- Strong proficiency in Python and Infrastructure as Code tools like Terraform or CloudFormation
- Solid understanding of distributed systems, networking, and GPU optimization for deep learning workloads
- Bonus: Experience deploying and serving large language models (LLMs) and generative AI applications in production
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search