Back to search
Bitus Labs Linkedin · Posted 6d ago

Data Engineer – AWS Lakehouse (Mandarin Required)

Irvine, California, United States

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Job description:

About the Role

We are looking for a mid-level Data Engineer to join our Data Platform team and take ownership of building and scaling our AWS-based data lakehouse. You will architect and deliver robust, production-grade data pipelines, work closely with data scientists, analytics engineers, and product teams, and set the technical direction for how data flows across the organization. This is a hands-on engineering role — you will write production code in Java and Python every day, while also contributing to platform design decisions, mentoring junior engineers, and driving best practices around data quality, reliability, and governance.

Key Responsibilities

Data Lakehouse Development

  • Build and extend medallion-architecture data lakehouse layers (Bronze /

Silver / Gold) on AWS S3 using the Apache Iceberg table format.

  • Develop and maintain high-throughput ETL/ELT pipelines using AWS

Glue, EMR (Spark), and Lambda.

  • Implement schema evolution, partitioning strategies, and compaction

processes for Iceberg tables to optimize storage and query performance.

  • Write production-quality pipeline code in Java and Python, following team

conventions for structure, testing, and maintainability.

Real-Time & Batch Streaming

  • Build and operate event-driven data pipelines using Amazon Kinesis Data

Streams, Kinesis Firehose, or Apache Kafka (MSK).

  • Implement exactly-once or at-least-once processing semantics for

streaming workloads using Apache Flink or Spark Structured Streaming on EMR.

AWS Platform Engineering

  • Contribute infrastructure-as-code definitions using AWS CDK or

Terraform for repeatable, auditable data platform deployments.

  • Monitor and help tune cost and performance across AWS services

including S3, Glue, Athena, Redshift Spectrum, EMR, Lambda, Step Functions, and

EventBridge.

  • Follow platform security standards in day-to-day work: IAM least-

privilege policies, KMS encryption, and VPC networking.

  • Maintain and extend CI/CD pipelines for data workloads using AWS

CodePipeline, GitHub Actions, or equivalent.

Data Quality & Governance

  • Add data quality checks using frameworks such as Great Expectations or

Deequ, and integrate validation steps into pipeline orchestration.

  • Help maintain and uphold data contracts between producing and

consuming systems.

  • Contribute to data cataloguing and lineage tracking using AWS Glue Data

Catalog or Apache Atlas.

Collaboration & Ways of Working

  • Partner with data scientists, ML engineers, and analysts to understand data

requirements and deliver performant, well-documented datasets.

  • Take an active part in code reviews, design discussions, and pair

programming, and support junior engineers where you can.

  • Document your pipeline designs and decisions, and contribute to the

internal engineering knowledge base.

Required Qualifications

Experience

  • 3+ years of professional data engineering experience, with at least 1–2

years on AWS cloud platforms.

  • Experience building and supporting production data pipelines, ideally on

large datasets with defined freshness or reliability SLAs.

  • Working knowledge of data lakehouse concepts — medallion pattern and

open table formats (Iceberg preferred; Delta Lake or Hudi acceptable).

Programming Languages

  • Java: Working proficiency in Java (8+) for Spark jobs and pipeline

components, with familiarity with Maven or Gradle build systems.

  • Python: Solid Python 3 skills for AWS Glue scripts, orchestration logic,

data quality checks, and automation tooling. Experience with pandas, PySpark, and

boto3. Strong skills in one language and a willingness to ramp up on the other is

acceptable.

AWS Core Services

  • Storage & Compute: S3, Glue (jobs, crawlers, Data Catalog), EMR

(Spark/Flink), Lambda, EC2.

  • Streaming: Kinesis Data Streams, Kinesis Firehose, or MSK (Managed

Kafka).

  • Orchestration: Step Functions, MWAA (Managed Airflow), or

EventBridge Scheduler.

  • Querying: Athena, Redshift, or Redshift Spectrum.
  • Security & Governance: Comfortable working within IAM, KMS, Secrets

Manager, and VPC setups.

  • DevOps: Exposure to AWS CDK or CloudFormation, and to CodePipeline

or equivalent CI/CD tools.

Data Processing Frameworks

  • Apache Spark (PySpark and/or Spark Java API) — distributed

transformations and a working grasp of performance tuning.

  • Apache Iceberg — reading and writing tables, time travel, and basic table

maintenance.

  • SQL — strong SQL for data transformation, including window functions,

CTEs, and query tuning.

  • Must be Chinese Mandarin fluent.

Preferred Qualifications

  • AWS Certified Data Engineer – Associate or AWS Certified Solutions

Architect certification.

  • Experience with dbt for SQL-based transformation layers on top of the

lakehouse.

  • Familiarity with ML platform integration: feature stores (SageMaker

Feature Store), model serving data needs, or MLflow experiment tracking.

  • Experience with real-time OLAP engines such as Apache Druid or

ClickHouse.

  • Experience with Lake Formation fine-grained access control, or with data

cataloguing and lineage tooling.

  • Exposure to data mesh or data product thinking — domain ownership and

data contracts.

Tech Stack at a Glance

Languages

Java/ Python 3

Cloud Platform

AWS (S3, Glue, EMR, Kinesis, Athena, Lambda, Step Functions, Lake Formation, CDK)

Processing

Apache Spark, Apache Flink, Spark Structured Streaming

Table Format

Apache Iceberg (primary), Delta Lake / Hudi (familiarity)

Streaming

Amazon Kinesis, MSK (Kafka), Kinesis Firehose

Orchestration

Apache Airflow (MWAA), AWS Step Functions

IaC & CI/CD

AWS CDK / Terraform, GitHub Actions / CodePipeline

Job Type: Full-time

Benefits:


  • 401(k)
  • 401(k) matching
  • Dental insurance
  • Health insurance
  • Life insurance
  • Paid time off
  • Parental leave
  • Retirement plan
  • Vision insurance


Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search