Data Engineer
Indexed description
Role Purpose:
The Lead ML Data Engineer is a senior technical leader responsible for enabling scalable, production-grade Data Science & Analytics (DSA) solutions within Coca-Cola Bottlers Japan Inc.'s (CCBJI) Vending Machines (VM) business unit. This role leads the development, optimization, and management of end-to-end ML and Analytics data workflows, ensuring reliable and efficient infrastructure for AI solutions like Assortment, Column Reallocation, and Placement.
The role works in close partnership with Data Scientists, Analytics Specialists, and IT to deliver high-impact, ML-ready datasets via best of breed tools. The role plays a key function in bridging business objectives with technical delivery by converting commercial data requirements into robust pipelines, reusable assets, and operational tooling.
The position also drives quality standards and data engineering practices within the team, ensuring model reproducibility, pipeline traceability, and integration with enterprise MLOps and governance frameworks.
Key Responsibilities:
Advanced ML Data Pipeline Development
- Design, develop, and maintain robust data pipelines to support ML model training, inference, and feature transformation workflows.
- Deliver performant and modular pipelines using Databricks (PySpark/Python) and Snowflake, aligned to architectural best practices.
- Ensure end-to-end ownership of data engineering from working with IT on raw data ingestion to developing model-ready feature layers.
- Implement CI/CD-ready transformation logic for reproducibility and handover to downstream components.
Feature Stores & Reusability Framework
- Architect and maintain a centralized, scalable feature stores that enables reuse across multiple DS use cases.
- Define feature documentation standards, naming conventions, and lifecycle management practices.
- Optimize joins, aggregations, and lookups to balance compute cost with accuracy and inference performance.
MLOps Integration & Model Lifecycle Engineering
- Work closely with the Data Scientists and Analytics Specialists to operationalize models through CI/CD pipelines (MLflow, GitHub Actions, Databricks Workflows).
- Design robust systems for retraining triggers, monitoring, and automated evaluation of production ML models.
- Implement failure recovery, alerting, and model rollback procedures in collaboration with IT and DevOps.
Engineering Excellence & Domain Leadership
- Serve as the go-to engineering authority for ML enablement within the DSA team.
- Lead technical design reviews, set code standards, and promote reusability and modularization across pipelines.
- Contribute internal tools, libraries, and utilities that improve engineering velocity and onboarding.
- Conduct informal mentoring and coaching for junior engineers and scientists working with data pipelines.
Agile Program Delivery
- Actively participate in agile sprint cycles, contributing to planning, estimation, retrospectives, and delivery metrics.
- Align with the Analytics Portfolio Manager, DS Manager, and Analytics Manager to prioritize deliverables and resolve cross-functional dependencies.
- Maintain and manage engineering backlog, surfacing technical debt or architectural decisions that require executive alignment.
Cross-Functional Collaboration
- Translate ambiguous business requirements into structured, scalable data solutions that accelerate DS and Analytics outcomes.
- Collaborate with the Analytics team to build curated views and pre-aggregated layers for Power BI or experimentation workflows.
- Partner with IT to ensure infrastructure provisioning, access control, and platform governance align with CCBJI’s enterprise standards.
Data Observability & Production Assurance
- Build and maintain monitoring systems to track pipeline performance, schema changes, and data freshness.
- Implement quality checks and exception handling to reduce operational risk and manual rework.
- Ensure SLAs are defined and met for model refreshes, data availability, and system uptime.
Key Outputs:
- Stable, scalable ML data pipelines that serve predictive models across VM business scenarios.
- Well-documented and reusable feature store logic shared across DS initiatives.
- Fully operationalized MLOps workflows supporting model retraining and deployment.
- Toolkits, templates, standards and frameworks adopted by team members for faster pipeline delivery.
- Reduction in time-to-ML model deployment time and increased delivery velocity.
- Measurable improvements in pipeline stability, data quality, and platform observability with operational dashboards
Performance Success Criteria (Examples):
- Launch production-grade ML pipelines for at least three strategic models within first 9–12 months.
- Reduce end-to-end model deployment cycle time by 25% through reusability and automation.
- Deliver a reusable feature store structure adopted by at least 3 DS initiatives.
- Implement and operationalize monitoring workflows covering pipeline reliability and data quality.
- Create internal engineering utilities or templates reused by at least two other team members.
- Maintain 98%+ reliability of ML workflows and data pipelines with documented support procedures.
Requirements:
Education:
- Bachelor’s or Master’s degree in Computer Science, Data Engineering, Information Systems, or a related quantitative discipline.
Experience:
- 6+ years in data engineering roles, with 2+ years focused on ML data workflows.
- Proven experience enabling ML model development and deployment through scalable engineering solutions.
- Hands-on experience with Snowflake and Databricks in production environments.
- Experience contributing to cross-functional data science or analytics programs.
- Experience designing feature stores and implementing feature pipelines in ML production environments.
- Working knowledge of CI/CD, data versioning, and DevOps practices for pipeline and model automation.
Technical Skills:
- Deep proficiency in Python and PySpark for data transformation and ML pipeline development in tools like Databricks.
- Advanced SQL capabilities for data wrangling, transformation, and pipeline optimization.
- Hands-on experience with MLflow, Git, Docker, and other DevOps tools.
- Working knowledge of model retraining, inference orchestration, and version control.
- Understanding of data governance, observability, and engineering compliance principles.
- Exposure to BI tools like Power BI for quick validation and prototyping.
Business & Communication Skills:
- Demonstrated ability to lead through expertise and technical credibility.
- Skilled in communicating complex architecture to business and non-technical partners.
- Works cross-functionally with DS, Analytics, and IT to drive aligned delivery.
- Embraces continuous improvement and knowledge-sharing mindset.
- Comfortable mentoring junior team members without formal people management.
- Driven by automation, reusability, and long-term scalability.
- Curious and proactive about learning new technologies and industry trends.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search