Lead Data Engineer (Databricks, PySpark & GCP)
Indexed description
Experience Level:
10 to 16 years of relevant IT experience
Key Responsibilities:
- Design, develop, test, and maintain scalable ETL data pipelines using Python, PySpark, Databricks & GCP / Azure.
- Architect the enterprise solutions with various technologies like GCP, Azure, Databricks, PySpark and Spark SQL.
- Work extensively on Google Cloud Platform (GCP) services such as:
- Dataflow for real-time and batch data processing
- Cloud Functions for lightweight serverless compute
- BigQuery for data warehousing and analytics
- Cloud Composer for orchestration of data workflows (on Apache Airflow)
- Google Cloud Storage (GCS) for managing data at scale
- IAM for access control and security
- Cloud Run for containerized applications
- Develop production-grade Databricks notebooks and workflows.
- Build data transformation pipelines using PySpark and Spark SQL.
- Implement Delta Lake architecture.
- Design Bronze, Silver, and Gold data layers using the Medallion Architecture.
- Implement Databricks Workflows/Jobs and dependency management.
- Tune Spark jobs for large-scale data processing.
- Optimize cluster configuration and compute utilization.
- Implement appropriate partitioning, caching, and file-size optimization strategies.
- Perform data ingestion from various sources and apply transformation and cleansing logic to ensure high-quality data delivery.
- Implement and enforce data quality checks, validation rules, and monitoring.
- Collaborate with data scientists, analysts, and other engineering teams to understand data needs and deliver efficient data solutions.
- Manage version control using GitHub and participate in CI/CD pipeline deployments for data projects.
- Write complex SQL queries for data extraction and validation from relational databases such as SQL Server, Oracle, or PostgreSQL.
- Document pipeline designs, data flow diagrams, and operational support procedures.
- 10+ years of hands-on experience in Python for backend or data engineering projects.
- Strong understanding and working experience with GCP cloud services (especially Dataflow, BigQuery, Cloud Functions, Cloud Composer, etc.).
- Working experience with Azure Data Factory (ADF), Azure Databricks, Azure Data Lake Storage Gen2 (ADLS).
- Solid understanding of data pipeline architecture, data integration, and transformation techniques.
- Experience in working with version control systems like GitHub and knowledge of CI/CD practices.
- Experience in Apache Spark, Kafka, Redis, Fast APIs, Airflow, GCP Composer DAGs.
- Strong experience in SQL with at least one enterprise database (SQL Server, Oracle, PostgreSQL, etc.).
- Experience with PySpark is required.
- Experience in data migrations from on-premise data sources to Cloud platforms.
- Experience with AWS services.
- Excellent problem-solving and analytical skills.
- Strong communication skills and ability to collaborate in a team environment.
- Bachelor's degree in Computer Science, a related field, or equivalent experience
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search