Back to search
Tesla Linkedin · Posted 1mo ago

Sr. Data Platform Engineer

Shanghai

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

The Role

We are building the next generation of enterprise data platforms: keeping production systems reliable while productizing query, processing, orchestration, and governance with solid engineering and mainstream AI capabilities.

You Will Bring

  • Platform engineering skills — production operations and component evolution for Spark / Flink, Airflow, Trino/Presto, and lakehouse table formats
  • Product sense — turn repetitive manual work into durable platform capabilities (APIs, CLIs, control planes, agents), not one-off scripts
  • AI fluency — use mainstream AI technologies (LLMs, agents, tool-calling, knowledge bases) to improve platform efficiency and experience, and take responsibility for generated outcomes

What You’ll Do

  • Own reliability, capacity, change management, and incident response for compute and orchestration planes (Spark / Flink / Airflow / Trino·Presto)
  • Maintain and evolve lakehouse processing and data-sync components, improving observability, release quality, and operational efficiency
  • Help modernize the query plane: cluster deployment, authorization and traffic control, reproducible environments, capacity planning, and slow-query remediation
  • Design platform capabilities with a product mindset: configuration contracts, approval boundaries, auditing, self-service onboarding, and closed-loop diagnostics
  • Explore and apply mainstream AI technologies to build next-generation data platform capabilities (intelligent troubleshooting, knowledge capture, natural-language interaction, automated operations), with human–AI co-governance rather than black-box automation
  • Partner with business, security, and infrastructure teams to improve platform experience and efficiency under compliance constraints

Must-have

  • 3+ years of experience in data platforms / big data engineering, with independent production on-call ownership
  • Compute & orchestration
  • Deep understanding of Spark / Flink, with hands-on production operations and troubleshooting
  • Deep understanding of Airflow, with hands-on scheduling operations and change management
  • Deep understanding of Trino / Presto, with hands-on cluster operations (concurrency/queues, authorization, slow queries, capacity)
  • Data lake / lakehouse
  • Deep understanding of Apache Hudi / Apache Iceberg (write modes, compaction/cleanup, time travel, and read/write paths with Spark/Trino)
  • Programming
  • Java / Scala required
  • Strong Linux troubleshooting and SQL skills
  • Product mindset: define platform capabilities from the user journey; prioritize reusable, measurable, operable outcomes over one-off firefighting

Nice-to-have

  • Vertica (or similar MPP / analytical warehouse) operations and performance tuning
  • Hands-on experience deploying and operating Spark / Flink / Trino on Kubernetes
  • Rust (strong plus for high-performance proxies, control planes, or platform tooling)
  • Python (orchestration ecosystem, automation, agent-side work)
  • Kafka / CDC experience (e.g., Debezium)
  • Metadata, lineage, or data-governance platform experience
  • Reverse proxy / gateway / quota-rate-limiting / observability experience (Prometheus, Grafana, logging platforms)
  • Agent / tool-calling / RAG / knowledge-base productization experience
Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search