Back to search
Zalo Linkedin · Posted 18d ago

Lead Data Engineer, Zalo (open for Senior)

Ho Chi Minh City

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Join the Audience Platform team at Zalo Group and architect the data backbone for one of Vietnam's largest digital ecosystems. With nearly 80 million Monthly Active Users and TBs of data flowing through our systems daily, we're looking for a Data Engineer to build robust pipelines that power large-scale data mining and give decision-makers seamless access to actionable insights.


Scope of work:

You won't just move data - you'll build the platforms that make data mining effortless. You will develop a centralized Feature Store and self-served, AI-backed automated analytics (AutoEDA) serving the Ads, Fintech, and VAS business lines. Your mission is to bridge massive datasets and business impact through thoroughly architected, highly optimized, rock-solid data infrastructure.


What you will do

  • Build and optimize data pipelines for diverse use cases (data mining, audience targeting, analytics) with strict compliance to data privacy regulations;
  • Engineer reliable pipelines to synchronize data across databases, powering batch + stream processing, large-scale analytics, and low-latency online serving;
  • Collaborate with Data Science, Policy, System Operation, and Research teams to understand their data needs and deliver solutions;
  • Establish solid design and engineering best practices for both technical and non-technical partners.


What you will need

Technical skills:

  • 5+ years hands-on in the Big Data ecosystem (Spark, Kafka, ClickHouse, Cassandra or similar);
  • Open table formats (Apache Iceberg / Delta / Hudi) with an open catalog (Unity / Polaris / Glue); storage & compute optimization (compression, partitioning, bucketing, ORC/Parquet, compaction, small-file handling);
  • Reliable pipelines syncing data across databases; CDC (Debezium), Kafka; unified batch + stream processing with exactly-once/idempotency, schema evolution, backfill;
  • Medallion Lakehouse; dimensional modeling / SCD / data vault; ELT (dbt) + semantic/metric layer;
  • Data contracts, testing, SLA/SLO, data observability & lineage;
  • Build the self-serve, AI-backed AutoEDA / analytics platform enabling stakeholders to mine data without engineering overhead (serving DS/ML — model training not required);
  • API & Platform: High-throughput APIs (gRPC, REST) on Kubernetes; latency & scalability optimization;
  • Airflow, Docker/containerization, CI/CD, observability (Prometheus, Grafana, OpenTelemetry), canary/blue-green deployment;
  • Strong Scala or Python for high-performance data processing; SQL;
  • Compliance with Vietnam's Personal Data Protection Law 2025 (Law 91/2025/QH15) & Decree 356/2025/NĐ-CP; GDPR for cross-border data;

Leadership & Stakeholder collaboration:

  • Mentor, code review, build engineering culture;
  • Collaborate with DS / BI / Product / Partner teams under a data-as-product mindset.

Nice to have:

  • Strong plus: HDFS / object-storage administration: backup, cleanup, archiving, capacity planning;
  • Centralized Feature Store (4,000+ features): point-in-time correctness, low-latency online serving, online/offline consistency, batch + real-time ingestion;
  • AutoML & MLOps: auto-training, evaluation, prediction at scale; model registry & serving; end-to-end ML pipeline understanding;
  • Fundamental knowledge of data science workflows and ML integration;
  • Experience with DMP/CDP and vector databases for AI/RAG use cases.


Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search