Big Data Engineer
Indexed description
Roles & Responsibilities
- Lead the design, development, and maintenance of robust, scalable data pipelines for ingestion, transformation, and processing of large datasets in an on-premises environment.
- Own architectural and design decisions for data solutions, evaluating trade-offs and defining technical standards for the team.
- Mentor, guide, and support other data engineers through code reviews, design reviews, technical coaching, and hands-on problem-solving.
- Build and optimize data workflows using Python, Spark, and the Hadoop ecosystem.
- Work extensively with Hadoop ecosystem components (Hive, HDFS, Impala) to manage and query large-scale data.
- Manage and optimize batch scheduling and job orchestration using enterprise schedulers such as CA7 or Control-M.
- Ensure data quality, integrity, and performance across data platforms.
- Collaborate with data analysts, data scientists, and business stakeholders to translate data requirements into sound technical designs.
- Troubleshoot and resolve complex issues in data pipelines and production environments, acting as an escalation point for the team.
- Champion best practices for coding standards, version control, testing, and documentation.
- Stay current with emerging technologies, particularly AI/ML capabilities, and identify opportunities to apply them to data engineering workflows.
Technical Skills
Must Have
- 7–10 years of overall experience in data engineering, with a proven track record in technical leadership (design ownership, mentoring, guiding development teams).
- Python – strong hands-on development experience building production-grade data solutions.
- Big Data / Hadoop ecosystem (Hadoop, Hive, Impala, HDFS) – deep, hands-on experience in on-premises environments.
- Apache Spark – solid experience developing and tuning large-scale distributed data processing jobs.
- Job scheduling / orchestration – hands-on experience with CA7 or Control-M (or comparable enterprise schedulers).
- Strong understanding of data structures, ETL processes, and SQL.
- Extensive experience with large-scale data processing and distributed systems.
- Demonstrated ability to make sound architecture/design decisions and to mentor and support other developers.
- Exposure to AI/ML concepts or tools, with a strong willingness to learn and grow in this space.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search