Data Engineer
Indexed description
- Design, develop, and maintain distributed data pipelines using Apache Spark (Core, SQL, Streaming) written in Scala.
- Build ingestion and transformation workflows across the Hadoop ecosystem — HDFS, Hive, YARN, MapReduce — for structured and semi-structured data at TB–PB scale.
- Tune and optimize Spark jobs: partitioning strategy, caching, broadcast joins, shuffle reduction, data skew handling, and executor/memory sizing.
- Implement real-time and near-real-time ingestion using Apache NiFi and/or Kafka.
- Embed data quality, reconciliation, and validation controls directly into pipelines.
- Author and optimize HiveQL and Spark SQL for curated and consumption layers.
- Automate orchestration and scheduling using Airflow, Oozie, or Control-M.
- Participate in code reviews, CI/CD automation, unit and integration testing, and production support.
- Troubleshoot job failures, SLA breaches, and performance regressions; drive root-cause analysis to permanent fixes.
- Document data flows, lineage, transformation logic, and operational runbooks.
- 6+ years of data engineering experience, with 4+ years hands-on Apache Spark development in Scala on production workloads.
- Strong Scala fundamentals — functional programming constructs, collections API, case classes, pattern matching, implicits, and error handling.
- Deep working knowledge of the Hadoop ecosystem: HDFS, Hive, YARN, HBase.
- Advanced SQL and data modeling skills across dimensional and big-data denormalized patterns.
- Demonstrated Spark performance tuning and debugging using the Spark UI, event logs, and physical execution plans.
- Proficiency with columnar and serialization formats — Parquet, ORC, Avro — including compression and partitioning trade-offs.
- Linux and shell scripting, Git, Maven or SBT, and Jenkins or equivalent CI/CD tooling.
- Ability to work independently in a distributed onshore–offshore delivery model.
- Kafka and Spark Structured Streaming for event-driven pipelines.
- Cloud data platform exposure — GCP (BigQuery, Dataproc), AWS EMR, or Azure Databricks.
- Telecom domain experience with clickstream, network, or geospatial/location data.
- Python or PySpark as a secondary development language.
- Data governance and security frameworks — Apache Ranger, Kerberos, PII masking and tokenization
- Apache NiFi flow design, configuration, and administration.
Qualifications
Bachelor
Range Of Year Experience-Min Year
15
Range Of Year Experience-Max Year
18
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search