Senior Data Engineer
Indexed description
Core Responsibilities
- Design and Maintain Data Pipelines: Build and maintain both batch and real-time data pipelines that are scalable, reliable, and production-ready.
- Integrate Diverse Data Sources: Ingest data from APIs, databases, files, and external sources using Apache Kafka and AWS services.
- Orchestrate Workflows: Develop and orchestrate workflows with Apache Airflow, including monitoring, retries, and alerting.
- Distributed Data Processing: Build and optimize distributed data processing solutions with Apache Spark, PySpark, Spark SQL, and AWS EMR.
- Manage Data Lake and Data Warehouse Environments: Administer Data Lake and Data Warehouse environments using Amazon S3 and Amazon Redshift.
- Analytical Data Modeling: Design analytical data models, including Star and Snowflake schemas, and prepare optimized datasets for Analytics and BI, particularly for Qlik.
- Engineering Best Practices: Deliver quality solutions in Python and SQL, following best practices for Git, code review, testing, and documentation.
- Data Quality, Security, and Governance: Ensure data quality, security, governance, and compliance, including AWS IAM, encryption, and access control.
- Production Support: Monitor production data pipelines and proactively address performance and operational issues.
- Proven experience as a Senior Data Engineer or in a similar role.
- Advanced proficiency in Python and SQL.
- Hands-on experience with Apache Spark / PySpark, Kafka, and Airflow.
- Strong AWS experience, particularly with S3, Redshift, and EMR.
- Experience with Data Lakes, Data Warehouses, ETL/ELT, and data modeling.
- Solid understanding of batch and streaming architectures.
- Experience with performance optimization, Git, testing, and production support.
This role requires familiarity with the following tools and technologies:
- Python and SQL: For data transformation, querying, and pipeline development.
- Spark / PySpark: For distributed data processing at scale.
- Kafka: To enable real-time data streaming and support event-driven architectures.
- Airflow: For scheduling, managing, and monitoring complex data workflows and ETL processes.
- AWS (EMR, S3, Redshift): For scalable data storage, processing, and warehousing.
- Git and Qlik: For version control and delivering analytics-ready datasets for BI.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search