Research Data Engineer
Indexed description
About The Role
As a Research Data Engineer, you will build the analysis library our engineering teams use to turn test data into decisions, and connect the data flows that feed it. Both build on existing data infrastructure. In this high-impact position, you will own the infrastructure and conventions that let data move reliably between teams and stay traceable from layout through to measurement, working directly with the people who use what you build.
Role and Responsibilities
Analysis library:
- Architect a Python library centered on pipeline orchestration for characterization data.
- Ensure abstraction between analysis logic and underlying file formats or storage schemas.
- Develop scalable solutions for local, remote, and distributed execution environments.
- Integrate logging and transformation lineage for dataframes with database export capabilities.
- Cater to both exploratory R&D and production-level quality assurance workflows.
- Provide documentation and support to facilitate technical adoption across internal teams.
- Build multi-site infrastructure for test data and metadata storage.
- Develop ingestion mechanisms for relational database metadata updates.
- Enable cross-team automation via metadata handshakes and metrology plan scripting.
- Standardize retrieval and storage conventions to ensure end-to-end data traceability.
- MSc in Computer Science, Engineering, or a scientific discipline (or equivalent practical experience).
- 5+ years of experience working with data in an R&D, research center, or instrument-heavy environment.
- Strong Python expertise with experience designing libraries and APIs.
- Proficiency with SQL, Git, and UNIX/bash.
- Experience building and running ETL/ELT pipelines with modern orchestrators (Prefect, Airflow, Dagster, or similar) in production.
- Familiarity with scientific data formats (HDF5, Zarr, Parquet) and structured data storage concepts (object storage, relational catalogs).
- Strong written and spoken English skills; ability to coordinate and unify workflows across independent teams.
- Openness to adopting new frameworks and state-of-the-art data tools.
- Domain expertise in integrated photonics or wafer-level fabrication processes.
- Practical experience with distributed computing frameworks (Ray, Dask, Spark) and high-volume data handling.
- Knowledge of data lineage systems and metadata cataloging solutions.
Practicalities
- Location: Lausanne, Switzerland (Onsite).
- Work percentage: 100%.
- Languages: English is our working language, other languages are a plus.
- Home office: Up to two days per week.
- Start date: As soon as possible.
Please submit your application through our career page.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search