Sr Data Engineer - Music DISCO, Music DISCO
Indexed description
This domain provides analytical support for the Consumer Product Tech org to make data driven decisions while launching new features and evaluating existing features with the end goal of improving the customer experience.
DISCO team enables repeatable, easy, in depth analysis of music customer behaviors. We reduce the cost in time and effort of analysis, data set building, model building, and user segmentation. Our goal is to empower all teams at Amazon Music to make data driven decisions and effectively measure their results by providing high quality, high availability data, and democratized data access through self-service tools.
If you love the challenges that come with big data then this role is for you. We collect billions of events a day, manage petabyte scale data on Redshift and S3, and develop data pipelines using Spark/Scala EMR, SQL based ETL, Airflow services.
We are looking for talented, enthusiastic, and detail-oriented Data Engineer, who knows how to take on big data challenges in an agile way. Duties include big data design and analysis, data modeling, and development, deployment, and operations of big data pipelines. You'll help build Amazon Music's most important data pipelines and data sets, and expand self-service data knowledge and capabilities through an Amazon Music data university.
DISCO team develops data specifically for a set of key business domains like personalization and marketing and provides and protects a robust self-service core data experience for all internal customers. We deal in AWS technologies like Redshift, S3, EMR, EC2, DynamoDB, Kinesis Firehose, and Lambda. Your team will manage the data exchange store (Data Lake) and EMR/Spark processing layer using Airflow as orchestrator. You'll build our data university and partner with Product, Marketing, BI, and ML teams to build new behavioural events, pipelines, datasets, models, and reporting to support their initiatives. You'll also continue to develop big data pipelines.
Key job responsibilities
You will work with Product Managers, Data scientists and other Data Engineers to help design, develop and deliver scalable data analytics platform and data pipeline solutions to support various Science, ML initiatives and at the scale and speed of Amazon Music. In addition, you will help design, develop, and deliver components for the analytics platform at the broader org level and streamline/automate workflows for the broader DISCO organization. You will serve as a leader for Data Engineers in DISCO MPT team.
A day in the life
- Collaborate with cross-functional teams, including data scientists, data scientists, business intelligence engineers, to design and architect a modern data analytics platform on AWS, utilizing the AWS Cloud Development Kit (CDK).
- Develop robust and scalable data pipelines using SQL/PySpark/Airflow to efficiently ingest, process, and transform large volumes of data from various sources into a structured format, ensuring data quality and integrity.
- Design and implement an efficient and scalable data warehousing solution on AWS, using appropriate NoSQL/SQL storage and database technologies for structured and unstructured data.
- Automate ETL/ELT processes to streamline data integration from diverse data sources and ensure the platform's reliability and efficiency.
- Create data models to support business intelligence, providing actionable insights and interactive reports to end-users.
- Enable advanced analytics and machine learning capabilities within the platform to derive predictive and prescriptive insights from the data through tools like EMR/SageMaker Notebooks
- Continuously monitor and optimize the performance of data pipelines, databases, and applications, ensuring low-latency data access for analytics and machine learning tasks.
- Implement robust security measures and ensure data compliance with internal requirements, industry standards, and regulations to safeguard sensitive information.
- Work closely with data scientists and business intelligence engineers to understand their requirements and collaborate on data-related projects.
- Create comprehensive technical documentation for the platform's architecture, data models, and APIs to facilitate knowledge sharing and maintainability.
Basic Qualifications
- 5+ years of data engineering experience
- Experience with data modeling, warehousing and building ETL pipelines
- Experience with SQL
- Experience in at least one modern scripting or programming language, such as Python, Java, Scala, or NodeJS
- Experience mentoring team members on best practices
- Experience with big data technologies such as: Hadoop, Hive, Spark, EMR
- Experience operating large data warehouses
Company - Servicios Comerciales Amazon Mexico S. de R.L. de C.V.
Job ID: A10478794
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search