Data Engineer – AWS Lakehouse (Mandarin Required)
Indexed description
Job description:
About the Role
We are looking for a mid-level Data Engineer to join our Data Platform team and take ownership of building and scaling our AWS-based data lakehouse. You will architect and deliver robust, production-grade data pipelines, work closely with data scientists, analytics engineers, and product teams, and set the technical direction for how data flows across the organization. This is a hands-on engineering role — you will write production code in Java and Python every day, while also contributing to platform design decisions, mentoring junior engineers, and driving best practices around data quality, reliability, and governance.
Key Responsibilities
Data Lakehouse Development
- Build and extend medallion-architecture data lakehouse layers (Bronze /
Silver / Gold) on AWS S3 using the Apache Iceberg table format.
- Develop and maintain high-throughput ETL/ELT pipelines using AWS
Glue, EMR (Spark), and Lambda.
- Implement schema evolution, partitioning strategies, and compaction
processes for Iceberg tables to optimize storage and query performance.
- Write production-quality pipeline code in Java and Python, following team
conventions for structure, testing, and maintainability.
Real-Time & Batch Streaming
- Build and operate event-driven data pipelines using Amazon Kinesis Data
Streams, Kinesis Firehose, or Apache Kafka (MSK).
- Implement exactly-once or at-least-once processing semantics for
streaming workloads using Apache Flink or Spark Structured Streaming on EMR.
AWS Platform Engineering
- Contribute infrastructure-as-code definitions using AWS CDK or
Terraform for repeatable, auditable data platform deployments.
- Monitor and help tune cost and performance across AWS services
including S3, Glue, Athena, Redshift Spectrum, EMR, Lambda, Step Functions, and
EventBridge.
- Follow platform security standards in day-to-day work: IAM least-
privilege policies, KMS encryption, and VPC networking.
- Maintain and extend CI/CD pipelines for data workloads using AWS
CodePipeline, GitHub Actions, or equivalent.
Data Quality & Governance
- Add data quality checks using frameworks such as Great Expectations or
Deequ, and integrate validation steps into pipeline orchestration.
- Help maintain and uphold data contracts between producing and
consuming systems.
- Contribute to data cataloguing and lineage tracking using AWS Glue Data
Catalog or Apache Atlas.
Collaboration & Ways of Working
- Partner with data scientists, ML engineers, and analysts to understand data
requirements and deliver performant, well-documented datasets.
- Take an active part in code reviews, design discussions, and pair
programming, and support junior engineers where you can.
- Document your pipeline designs and decisions, and contribute to the
internal engineering knowledge base.
Required Qualifications
Experience
- 3+ years of professional data engineering experience, with at least 1–2
years on AWS cloud platforms.
- Experience building and supporting production data pipelines, ideally on
large datasets with defined freshness or reliability SLAs.
- Working knowledge of data lakehouse concepts — medallion pattern and
open table formats (Iceberg preferred; Delta Lake or Hudi acceptable).
Programming Languages
- Java: Working proficiency in Java (8+) for Spark jobs and pipeline
components, with familiarity with Maven or Gradle build systems.
- Python: Solid Python 3 skills for AWS Glue scripts, orchestration logic,
data quality checks, and automation tooling. Experience with pandas, PySpark, and
boto3. Strong skills in one language and a willingness to ramp up on the other is
acceptable.
AWS Core Services
- Storage & Compute: S3, Glue (jobs, crawlers, Data Catalog), EMR
(Spark/Flink), Lambda, EC2.
- Streaming: Kinesis Data Streams, Kinesis Firehose, or MSK (Managed
Kafka).
- Orchestration: Step Functions, MWAA (Managed Airflow), or
EventBridge Scheduler.
- Querying: Athena, Redshift, or Redshift Spectrum.
- Security & Governance: Comfortable working within IAM, KMS, Secrets
Manager, and VPC setups.
- DevOps: Exposure to AWS CDK or CloudFormation, and to CodePipeline
or equivalent CI/CD tools.
Data Processing Frameworks
- Apache Spark (PySpark and/or Spark Java API) — distributed
transformations and a working grasp of performance tuning.
- Apache Iceberg — reading and writing tables, time travel, and basic table
maintenance.
- SQL — strong SQL for data transformation, including window functions,
CTEs, and query tuning.
- Must be Chinese Mandarin fluent.
Preferred Qualifications
- AWS Certified Data Engineer – Associate or AWS Certified Solutions
Architect certification.
- Experience with dbt for SQL-based transformation layers on top of the
lakehouse.
- Familiarity with ML platform integration: feature stores (SageMaker
Feature Store), model serving data needs, or MLflow experiment tracking.
- Experience with real-time OLAP engines such as Apache Druid or
ClickHouse.
- Experience with Lake Formation fine-grained access control, or with data
cataloguing and lineage tooling.
- Exposure to data mesh or data product thinking — domain ownership and
data contracts.
Tech Stack at a Glance
Languages
Java/ Python 3
Cloud Platform
AWS (S3, Glue, EMR, Kinesis, Athena, Lambda, Step Functions, Lake Formation, CDK)
Processing
Apache Spark, Apache Flink, Spark Structured Streaming
Table Format
Apache Iceberg (primary), Delta Lake / Hudi (familiarity)
Streaming
Amazon Kinesis, MSK (Kafka), Kinesis Firehose
Orchestration
Apache Airflow (MWAA), AWS Step Functions
IaC & CI/CD
AWS CDK / Terraform, GitHub Actions / CodePipeline
Job Type: Full-time
Benefits:
- 401(k)
- 401(k) matching
- Dental insurance
- Health insurance
- Life insurance
- Paid time off
- Parental leave
- Retirement plan
- Vision insurance
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search