Senior Data Platform / Data Engineer
Indexed description
About The Role
We are looking for a Senior Data Platform / Data Engineer to join our ML Platform team and help build and scale the data infrastructure that powers our AI products in dentistry. Our platform supports the full AI development lifecycle, from raw data ingestion and annotation workflows to dataset versioning and model training pipelines. You will work closely with Machine Learning Researchers (MLRs), MLOps engineers, and product teams to ensure our data infrastructure is reliable, scalable, and easy to use.
A key focus of the role is improving our Data Lakehouse (DLH) and dataset management workflows, including dataset versioning (DVC) and improving how data is prepared, extracted, and consumed across research and production systems.
What You Will Work On
You will play a key role in shaping the next generation of our data platform.
Typical Responsibilities Include
Data platform ownership
- Design and evolve the Data Lakehouse (DLH) architecture used across our ML teams.
- Improve the reliability and structure of data ingestion, extraction, and transformation pipelines.
- Ensure datasets used for training and evaluation are consistent, reproducible, and well documented.
- Improve workflows for dataset versioning and reproducibility using tools such as DVC.
- Design solutions for managing multiple versions of datasets and annotations across experiments and models.
- Improve the ability for researchers to retrieve the correct dataset versions reliably.
- Build and maintain scalable data pipelines in Python.
- Improve metadata management, dataset validation, and data quality monitoring.
- Optimize data workflows across AWS-based infrastructure.
- Work closely with ML researchers and ML engineers to understand their data needs.
- Support research workflows with reliable and efficient data access patterns.
- Help translate research requirements into robust platform capabilities.
- Implement practices for data quality, reproducibility, and traceability across the ML lifecycle.
- Ensure our data infrastructure meets the requirements of regulated AI development.
- Strong Python engineering skills
- Experience building data pipelines or data platforms
- Experience working with AWS
- Experience working with large datasets used in ML workflows
- Strong software engineering practices (testing, CI/CD, documentation)
- Experience collaborating with ML teams or working in AI environments
- Experience with dataset versioning tools such as DVC
- Experience with Kubernetes
- Experience with data lakehouse architectures
- Experience working with annotation pipelines or ML training datasets
- Experience with PostgreSQL, Metabase, or similar data tooling
- Experience working in regulated environments (medical / healthcare AI)
- AWS
- Python
- Kubernetes
- PostgreSQL
- Metabase
- DVC for dataset versioning
- Internal Data Lakehouse infrastructure
Employment Type: Full Time
Alternative Locations: Spain : Madrid || Poland : Gdansk || Poland : Warsaw || Poland : Wroclaw
Travel Percentage: 0 - 10%
Requisition ID: 20071
20071
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search