Lead Data Engineer
Indexed description
Key Responsibilities
Act as SME on GCP data engineering patterns — Dataproc (Spark/Hadoop) jobs, Cloud Composer (Airflow) DAGs, and BigQuery data models used in the target architecture
Guide the migration team in re-platforming legacy ETL/data warehouse jobs onto GCP, including orchestration design, job scheduling, and dependency mapping in Composer
Design and build Dataproc pipelines to process and transform migrated data, ensuring parity with legacy business logic
Collaborate with data engineers and migration leads to troubleshoot discrepancies during testing and validation of migrated pipelines
Optimize BigQuery performance (partitioning, clustering, query design) and Dataproc cluster/job configuration for cost and efficiency
Required Skills
Fluency in both German and English
Hands-on experience with core GCP data services — BigQuery, Dataproc, Cloud Composer (Airflow), Cloud Storage
Strong Python and/or PySpark skills for building and maintaining data pipelines
Solid understanding of Airflow DAG design, scheduling, and dependency management
Experience with SQL and data modeling for analytical workloads
Preferred / Nice-to-Have
Prior experience on a legacy-to-cloud data migration project
Familiarity with legacy ETL tools (Ab Initio) or on-prem data warehouses for migration context
Experience with CI/CD for data pipelines (Cloud Build, Terraform)
Exposure to job scheduling tools such as UC4, and relational databases (Oracle, SQL Server)
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search