Data Engineer
Indexed description
## About the Role
We are looking for a highly skilled Data Engineer with 4–6 years of experience in designing, building, and maintaining scalable data platforms and data pipelines. The ideal candidate should possess strong expertise in cloud-based data engineering, ETL/ELT development, data warehousing, and big data technologies. The candidate will work closely with Data Scientists, Analysts, BI teams, and business stakeholders to enable reliable and efficient data solutions.
## Key Responsibilities
* Design, develop, and maintain scalable ETL/ELT pipelines using Python, PySpark, SQL, Pandas, and DBT.
* Build and optimize batch and real-time data processing workflows using Apache Airflow and Databricks.
* Develop and manage cloud-based data lakes and data warehouses using AWS services such as S3, EMR, Redshift, Lambda, and RDS.
* Implement and maintain Snowflake data warehouse solutions, including Snowpipe, external stages, and data-sharing capabilities.
* Develop robust data ingestion frameworks for structured and semi-structured data from multiple sources.
* Design and implement event-driven architectures using AWS Lambda, Kafka, S3 Event Notifications, and other cloud-native services.
* Collaborate with cross-functional teams including Data Science, Analytics, BI, and Product teams to support data requirements.
* Build and expose APIs for data integration and data-sharing use cases.
* Ensure data quality, governance, privacy, and compliance through validation, masking, and monitoring frameworks.
* Implement CI/CD pipelines for data engineering projects using GitHub Actions, GitLab, JFrog, and automated testing frameworks.
* Optimize data pipelines for performance, scalability, and cost efficiency.
* Create operational dashboards and reporting solutions using Tableau or similar BI tools.
## Required Skills & Qualifications
### Technical Skills
* Strong programming skills in Python and SQL.
* Hands-on experience with PySpark, Pandas, and large-scale data processing.
* Expertise in Databricks, DBT, and Apache Airflow.
* Experience with Snowflake Data Warehouse and Snowpipe.
* Strong understanding of AWS ecosystem including:
* S3
* EMR
* Redshift
* Lambda
* EC2
* RDS
* Experience working with Apache Kafka and event-driven data architectures.
* Hands-on experience with PostgreSQL, MongoDB, and relational databases.
* Knowledge of Docker and containerized deployments.
* Experience with REST APIs and data integration frameworks.
* Familiarity with GitHub, GitLab, CI/CD pipelines, and automated testing practices.
* Understanding of Data Lake and Data Warehouse architectures.
### Preferred Qualifications
* Experience in Healthcare, Life Sciences, Retail, or eCommerce domains.
* Knowledge of data governance, security, and compliance best practices.
* Exposure to cloud certifications (AWS, OCI, Azure, or GCP).
* Experience working in Agile/Scrum environments.
## Education
* Bachelor's Degree in Computer Science, Information Technology, Engineering, or a related field.
## Nice to Have
* Experience with DuckDB.
* Exposure to Generative AI, Machine Learning, or MLOps platforms.
* Experience designing enterprise-scale data platforms and modern lakehouse architectures.
## What Success Looks Like
* Deliver highly reliable and scalable data pipelines.
* Ensure high-quality and trusted data for analytics and business decision-making.
* Improve data processing efficiency and reduce operational overhead.
* Contribute to the evolution of the organization's modern data platform strategy.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search