MLOps Engineer
Indexed description
The team supports data scientists and ML engineers working on high-volume AI products, with a focus on reproducibility, deployment velocity, cost control, and operational reliability. This role is based in Dallas, TX and is remote.
Key Responsibilities
- Build and maintain CI/CD and continuous training pipelines for ML models using Python, Docker, Kubernetes, and GitHub Actions or similar tooling
- Automate model packaging, validation, versioning, deployment, and rollback workflows across development, staging, and production environments
- Provision and manage cloud infrastructure with Terraform across AWS, GCP, or Azure, including compute, storage, networking, and managed ML services
- Operate scalable model serving systems using Kubernetes, KServe, Seldon, Ray Serve, or equivalent frameworks, optimizing latency and resource utilization
- Implement observability for model and platform health using Prometheus, Grafana, OpenTelemetry, and centralized logging
- Establish monitoring for data drift, model performance, feature quality, pipeline failures, and infrastructure regressions with actionable alerting
- Partner with data scientists, ML engineers, and software teams to define deployment standards, improve developer workflows, and resolve production incidents
- 3–8 years of experience in MLOps, platform engineering, DevOps, or software engineering supporting machine learning systems in production
- Strong Python and Linux skills, with practical experience building automation, APIs, and infrastructure tooling
- Hands-on experience with Docker, Kubernetes, Helm, and infrastructure as code using Terraform or an equivalent tool
- Experience implementing ML lifecycle workflows with platforms or tools such as MLflow, Kubeflow, SageMaker, Vertex AI, Azure ML, or Databricks
- Proficiency with cloud infrastructure and CI/CD systems, including networking, IAM, secrets management, artifact registries, and deployment automation
- Bachelor’s degree in computer science, engineering, mathematics, or a related technical field, or equivalent professional experience
- Bonus: Experience with feature stores, distributed training, GPU scheduling, Spark, Argo Workflows, model governance, or production LLM serving
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search