Data Engineer
Indexed description
Role Overview Seeking a PySpark Data Engineer with at least 2 years of hands-on experience in developing scalable data pipelines, ETL processes, and big data solutions using PySpark and Python. The ideal candidate will have experience working with large datasets, distributed data processing frameworks, and modern cloud-based data platforms. Key Responsibilities Develop and maintain ETL/ELT pipelines using PySpark. Build and optimize data processing workflows for large-scale datasets. Implement data transformation, validation, and cleansing logic. Write efficient Spark SQL and Python code for data engineering solutions. Troubleshoot and optimize Spark jobs for performance and scalability. Collaborate with architects, analysts, and business stakeholders to deliver data solutions. Support data quality, governance, and production operations. Required Skills 2+ years of experience with PySpark. Strong proficiency in Python and SQL. Experience with data engineering, ETL/ELT development, and data transformation. Good understanding of Apache Spark architecture and distributed computing concepts. Experience working with large datasets and performance optimization. Familiarity with data warehousing and data modeling concepts. Experience with Git and software development best practices. Preferred Skills Databricks. Snowflake, BigQuery, Synapse, Redshift, or other cloud data warehouses. Azure, AWS, or GCP. Kafka or streaming technologies. Airflow or other orchestration tools. Delta Lake and Lakehouse architecture. Qualifications Bachelor's degree in Computer Science, Information Technology, Engineering, or related field. Relevant certifications such as Databricks, Azure Data Engineer, AWS Data Analytics, or Google Professional Data Engineer are a plus. Minimum 5 year(s) of experience is required
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search