Data Engineer
Indexed description
🚀 Data Engineer | Azure Data & AI, Databricks, PySpark | AI Incubator Platform
🌍 Remote | 📄 B2B Contract |
We are looking for a detail-oriented Data Engineer with strong Data Science and Machine Learning lifecycle awareness to join our client's team—an in-house technology consultancy at a global consulting firm. You will design, build, and operate scalable, production-ready data pipelines for advanced analytics and NLP/NLU models, scouting and prototyping cutting-edge solutions across a 3–5 year horizon.
What You’ll Be Doing:
- Building, hardening, and maintaining reliable data ingestion and ETL/ELT pipelines for enterprise data platforms.
- Developing large-scale data processing and transformation workflows using Databricks and PySpark.
- Operating and integrating Azure data and AI infrastructure, including Blob Storage, compute resources, MLflow, and Azure AI/ML services.
- Preparing high-quality datasets, feature pipelines, and data architectures to support downstream NLP, NLU, and generative AI use cases.
- Implementing CI/CD and MLOps practices, including automated testing, deployment pipelines, observability, monitoring, and quality promotion gates.
- Taking end-to-end ownership of data quality, lineage tracking, and operational handovers for production-ready data products.
What We’re Looking For:
- Proven commercial experience in Data Engineering, large-scale data transformation, and production pipeline deployment.
- Strong Data Science and ML lifecycle awareness with the ability to bridge data engineering workflows with model experimentation and production deployment.
- Consultant mindset: proactive problem-solving, architectural ownership, and clear communication with cross-functional technical teams.
- Experience with pipeline observability, schema design, data quality validation, and lineage tracking.
Must-Haves:
- Strong programming and data processing skills in Python and PySpark.
- Hands-on expertise with Databricks for distributed data engineering and scalable workflow orchestration.
- Practical experience with the Azure ecosystem (Azure Blob Storage, Azure SQL/Databases, compute resources, MLflow, Azure AI/ML services).
- Solid track record supporting NLP/NLU workflows and preparing structured datasets for machine learning applications.
- Strong English communication skills for daily international collaboration.
Location & Working Setup:
- 100% Remote working model (Candidates must be living and working within the EU).
- US Hours overlap needed: Standard work hours by 6:00 PM CET (Typically 10:00 AM – 6:00 PM or 11:00 AM – 7:00/8:00 PM CET).
- Availability for occasional in-person team workshops/gatherings (approx. 1x per quarter in Prague).
Ready to build the data foundation for next-generation AI platforms? Apply today or send your CV directly to discuss the details!
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search