MLOps Engineer
Indexed description
The team works closely with machine learning engineers and data scientists to build robust MLOps practices, automated CI/CD pipelines, and comprehensive monitoring systems.
Key Responsibilities
- Design, build, and maintain scalable MLOps infrastructure and pipelines using Kubernetes, Docker, and Terraform on cloud platforms
- Implement automated CI/CD pipelines for model training, testing, and deployment to ensure rapid and reliable releases
- Manage and optimize model serving infrastructure using Triton, TorchServe, or vLLM to achieve strict latency and throughput SLAs
- Set up end-to-end monitoring and observability systems for data drift, concept drift, and resource utilization using Prometheus and Grafana
- Collaborate with data engineers to enforce feature store integrity and ensure seamless parity between training and inference data
- Establish security, governance, and cost-optimization practices for all cloud-hosted machine learning workloads
- 3–7 years of experience in MLOps, DevOps, or machine learning engineering with a heavy focus on production infrastructure
- Strong proficiency in Kubernetes, containerization, and infrastructure-as-code tools like Terraform or CloudFormation
- Hands-on experience with modern model serving frameworks and building automated model CI/CD pipelines
- Solid understanding of cloud platforms such as AWS, GCP, or Azure, including IAM, networking, and cluster management
- Strong software engineering fundamentals in Python and Bash scripting for automation and tooling
- Bonus: Experience managing LLM infrastructure, vector databases, or high-throughput real-time inference systems at scale
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search