Tekonika - Data Engineer
Indexed description
Work from Office opportunity - Client Interaction Involved.
Key Responsibilities
- Architect and deploy scalable data pipelines to ensure seamless data ingestion and transformation across diverse enterprise environments.
- Implement and manage modern lakehouse architectures to provide a unified, performant interface for both batch and real-time analytics.
- Develop and optimize streaming applications using Apache Flink and Kafka to support low-latency data processing requirements.
- Automate data transformation workflows using dbt to ensure high-quality, consistent, and well documented data models for downstream consumption.
- Manage the lifecycle of data stored in Apache Iceberg tables to improve query performance and data reliability for analytical workloads.
- Collaborate with cloud infrastructure teams to provision and maintain secure, cost-effective environments that support high-volume data processing.
- Proficient in Python and SQL.
- Expert in building large-scale data pipelines and lakehouse architectures using PySpark, DBT, Apache Kafka / Apache Flink, and Apache Iceberg.
- Strong experience with AWS, including cost optimization (e.g., reducing EMR costs) and building cloud-agnostic MLOps platforms.
- Hands-on expertise with Apache Airflow for automating and scaling production pipelines.
- Experience with Snowflake (Snow pipes, DBT models), Presto, and Ceph RADOS.
- Skilled in deploying NLP/ML models (such as sentence transformers and collaborative filtering), building end-to-end MLOps solutions, and feature engineering.
- Proficient in creating insights through Tableau and Python-based Dash.
- Experience evaluating and implementing tools like Immuta, Collibra, Atlan, and CastorDoc for cataloging and security.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search