Senior Data Engineer, Modern Data Stack and AI Pipelines
Indexed description
What You Will Do
- Design and build scalable, production-grade data pipelines for both batch and real-time workloads using Apache Spark, Kafka, and cloud-native tools
- Own the ELT/ETL architecture using dbt, Airflow (or Prefect/Dagster), and cloud data warehouses (Snowflake, BigQuery, Databricks Delta Lake)
- Build and maintain feature stores for ML teams, Feast, Tecton, or custom, ensuring consistent feature availability between training and serving
- Implement data quality frameworks, Great Expectations, Monte Carlo, or Soda, with automated monitoring and alerting
- Design data lakehouse architectures using Delta Lake or Apache Iceberg, time travel, ACID transactions, schema evolution
- Build real-time data streaming pipelines using Kafka, Kinesis, or Pub/Sub for low-latency ML inference serving
- Implement data governance and lineage tracking, dbt docs, OpenLineage, Apache Atlas, or Collibra
- Collaborate closely with ML engineers on feature engineering, training data pipelines, and model monitoring data infrastructure
- Define and enforce data engineering best practices, testing, documentation, version control, CI/CD for data pipelines
- 5+ years of data engineering experience with production pipeline ownership
- Strong proficiency in Python and SQL for data transformation and pipeline development
- Modern data stack expertise, dbt, Airflow/Prefect, and at least one cloud data warehouse (Snowflake, BigQuery, Databricks)
- Apache Spark experience for large-scale data processing
- Real-time streaming experience with Kafka or equivalent
- Cloud data platform experience, AWS (Glue, S3, Redshift), Azure (Data Factory, ADLS, Synapse), or GCP (Dataflow, BigQuery)
- Data quality and observability tooling experience
- Strong understanding of data modelling, dimensional modelling, Data Vault, or medallion architecture
- Feature store implementation experience for ML pipelines
- Delta Lake or Apache Iceberg expertise
- Data governance and lineage tooling experience
- Experience with LLM data pipelines, chunking, embedding, vector database ingestion
- dbt certification or Databricks certifications
- Data engineering ownership on AI-critical infrastructure, your pipelines power the models
- Modern toolstack with no legacy constraints, dbt, Spark, cloud-native from Day 1
- Remote-first flexibility with strong engineering peer community
- Competitive compensation reflecting acute market scarcity for this profile
dbtSparkKafkaSnowflakeDatabricksAirflowPythonAI Pipelines
Apply
Apply for Senior Data Engineer, Modern Data Stack and AI Pipelines
Complete the form below. Our team reviews every application personally, no automated filtering, no keyword matching. We will be in touch within two business days.
Full Name
Phone
+91 +1 +44 +971 +65 +61 +49 +353
LinkedIn Profile URL
GitHub Profile URL
Portfolio / Website URL
Optional, but sharing any of these helps us get to know your work faster.
Current Location
Open to Relocation?
Current CTC
Expected CTC
Years of Relevant Experience
Cover Note / Why This Role
Upload Resume
Drag and drop or browse to upload
PDF, DOC, or DOCX, max 5MB
Your application is reviewed by a domain expert, not an ATS. We do not share your details without your consent. See our Privacy Policy.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search