Databricks MLOps Engineer
Indexed description
This is a hands-on engineering role. You'll be writing infrastructure-as-code, building data pipelines, integrating LLMs into production systems, and collaborating closely with AI and data engineers to keep the platform performing as the workloads grow.
What You'll Do
- Design, implement, and maintain a cloud-native platform to support AI and data workloads, with a focus on Databricks and AWS Bedrock.
- Build and manage scalable data pipelines to ingest, transform, and serve data for ML and analytics use cases.
- Develop infrastructure-as-code using CloudFormation and AWS CDK to ensure repeatable, secure deployments.
- Integrate AI models and LLMs into production systems, including RAG architectures and model serving workflows.
- Drive observability best practices — monitoring, alerting, and logging across AI platforms.
- Collaborate with AI engineers, data engineers, and platform teams to improve performance, reliability, and cost-efficiency of models in production.
- Contribute to the design and evolution of the AI platform to support new ML frameworks, workflows, and data types.
- Stay current with emerging tooling and recommend improvements to architecture and operations.
- 7+ years of professional experience in software and infrastructure engineering.
- Extensive experience building and maintaining AI/ML infrastructure in production — model deployment, lifecycle management, and serving workflows.
- Production-level experience with Databricks and MLflow — model registration, versioning, asset bundles, and serving.
- Strong knowledge of AWS and infrastructure-as-code, with AWS CDK experience specifically.
- Expert-level coding skills in both Python and TypeScript — building robust APIs and backend services.
- Strong understanding of containerization with Docker and hands-on CI/CD pipeline experience.
- Proven ability to design reliable, secure, and scalable infrastructure for real-time and batch ML workloads.
- Strong communication and collaboration skills — able to articulate technical decisions clearly to both engineering and non-technical stakeholders.
- Familiarity with DSPy or similar LLM orchestration frameworks for programmatic prompt pipelines.
- Experience with LLM cost monitoring, latency optimization, and usage analytics in production.
- Knowledge of vector databases and embedding stores — OpenSearch or similar — to support semantic search and RAG architectures.
- Experience with ECS or other container orchestration tools.
- Shape real-world AI-driven projects across key industries, working with clients from startup innovation to enterprise transformation.
- Be part of a global team with equal opportunities for collaboration across continents and cultures.
- Thrive in an inclusive environment that prioritizes continuous learning, innovation, and ethical AI standards.
Solvd is an equal opportunity employer.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search