Lead Data Engineer
Indexed description
The Lead Data Engineer is crucial to D&A USB’s Eligibility team for creating and optimizing data architecture, solutions, and operations, and for ensuring alignment with data management and governance framework. The Lead Data Engineer serves as a big data development expert within the D&A Eligibility organization.
This position is responsible for leading, architecting, and building ETL, data warehousing, and reusable components using cutting-edge big data and cloud technologies. The resource will collaborate with the architect, business systems analyst, technical leads, project managers, and business/operations teams in building data enablement solutions across different LOBs and use cases.
Key Responsibilities
- Design and solution end-to-end data architecture for data hubs/data products, all the way from source systems to consumption.
- Oversee the design and management of data solutions to ensure data is stored, processed, curated, and utilized effectively.
- Own and build a reusable data pipeline utilizing Big Data and Azure/Databricks for Eligibility Program data products.
- Ingesting huge volumes of data from various platforms for Analytics needs and writing high-performance, reliable, and maintainable ETL code.
- Leadership: Lead and mentor a team of data engineers, ensuring the efficient flow of data within the organization with the defined processes and tools.
- Collect, store, process, and analyze large datasets to build and implement extract, transfer, load (ETL) processes.
- Develop reusable frameworks to reduce the development effort involved, thereby ensuring cost savings for the projects.
- Utilizing CI/CD Pipelines: Utilize and enhance CI/CD practices to automate the delivery of data solutions, ensuring reliability and scalability based on the defined tools.
- Utilize Cloud technologies (Azure Databricks) to enable data product solutions.
- Develop quality code through performance optimizations in place right at the development stage.
- Appetite to learn new technologies and be ready to work on new cutting-edge cloud technologies.
- Partner with Tech, Business, BI, and Data Science teams to create reusable data products.
- Work with a team spread across the globe in driving the delivery of projects and recommend development and performance improvements.
- Track and report on KPIs for solution delivery and data quality.
- Communicate and present use cases, solutions, and impact to business stakeholders and mid/senior management.
- Optimize reusable frameworks, Spark jobs for performance and cost efficiency in large-scale environments.
- Ability to interact with business stakeholders in getting the requirements and implementing solutions.
- Analyze the data in depth, using SQLs and other exploratory tools against various platforms such as Bigdata, Oracle, SQL Server, Databricks and others.
- Work with IT, business, and architects to develop and design requirements to formulate technical design.
- 8+ years of Data Solutions, development, and delivery experience with 4+ years of recent experience in Azure/Databricks environments.
- Proficiency and extensive experience with SQL, Spark &/or Scala/Python and performance tuning.
- Hands-on expertise in: Big Data (ex: Hive and HBase), Azure Databricks, Azure Functions, Cosmos DB and/or Data Factory experience is a MUST.
- Strong experience in building/designing Data warehouses, data stores for analytics consumption on Cloud (real time as well as batch use cases).
- Design, build, and deploy robust data ingestion and curation pipelines utilizing cloud-based data platforms such as Azure Data Factory, Apache Spark (Scala or Python), Azure Databricks, and Delta Lake.
- Good scripting experience, primarily on shell/bash/ PowerShell.
- Strong SQL knowledge and data analysis skills for data anomaly detection and data quality assurance.
- Experience and familiarity implementing data governance and data quality using the enterprise toolset.
- Skilled in drafting functional, technical requirements, creating high-level design documents, and data-flow diagrams, etc.
- Expertise in writing validation scripts to validate the data, data integrations and ETL transformations.
- Very good problem solver and excellent communication skills - both written and verbal.
- Expertise in Python and experience writing Azure functions using Python/Node.js.
- Databricks certifications and/ or Microsoft Azure Certifications
- Experience using Event Hub for data integrations.
- Eagerness to learn new technologies on the fly and ship to production.
- Hive database management and Performance tuning - Partitioning / Bucketing.
- Ability to interact with senior leadership teams in IT and business.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search