Back to search
Infinite Computer Solutions Linkedin · Posted 1mo ago

Data Engineer

Texas, United States

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Job Description

Role Summary We are seeking an experienced Big Data Engineer to design, build, and optimize large-scale batch and streaming data pipelines on the Hadoop ecosystem using Apache Spark and Scala. The role supports high-volume ingestion, transformation, and enrichment of clickstream, network, and location datasets, working closely with data architects, platform engineering, and downstream analytics teams. This is a hands-on engineering role with ownership of pipeline performance, reliability, and data quality in production. Key Responsibilities

  • Design, develop, and maintain distributed data pipelines using Apache Spark (Core, SQL, Streaming) written in Scala.
  • Build ingestion and transformation workflows across the Hadoop ecosystem — HDFS, Hive, YARN, MapReduce — for structured and semi-structured data at TB–PB scale.
  • Tune and optimize Spark jobs: partitioning strategy, caching, broadcast joins, shuffle reduction, data skew handling, and executor/memory sizing.
  • Implement real-time and near-real-time ingestion using Apache NiFi and/or Kafka.
  • Embed data quality, reconciliation, and validation controls directly into pipelines.
  • Author and optimize HiveQL and Spark SQL for curated and consumption layers.
  • Automate orchestration and scheduling using Airflow, Oozie, or Control-M.
  • Participate in code reviews, CI/CD automation, unit and integration testing, and production support.
  • Troubleshoot job failures, SLA breaches, and performance regressions; drive root-cause analysis to permanent fixes.
  • Document data flows, lineage, transformation logic, and operational runbooks.

Required Qualifications

  • 6+ years of data engineering experience, with 4+ years hands-on Apache Spark development in Scala on production workloads.
  • Strong Scala fundamentals — functional programming constructs, collections API, case classes, pattern matching, implicits, and error handling.
  • Deep working knowledge of the Hadoop ecosystem: HDFS, Hive, YARN, HBase.
  • Advanced SQL and data modeling skills across dimensional and big-data denormalized patterns.
  • Demonstrated Spark performance tuning and debugging using the Spark UI, event logs, and physical execution plans.
  • Proficiency with columnar and serialization formats — Parquet, ORC, Avro — including compression and partitioning trade-offs.
  • Linux and shell scripting, Git, Maven or SBT, and Jenkins or equivalent CI/CD tooling.
  • Ability to work independently in a distributed onshore–offshore delivery model.

Preferred Qualifications

  • Kafka and Spark Structured Streaming for event-driven pipelines.
  • Cloud data platform exposure — GCP (BigQuery, Dataproc), AWS EMR, or Azure Databricks.
  • Telecom domain experience with clickstream, network, or geospatial/location data.
  • Python or PySpark as a secondary development language.
  • Data governance and security frameworks — Apache Ranger, Kerberos, PII masking and tokenization

Nice to Have:

  • Apache NiFi flow design, configuration, and administration.

Education Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline — or equivalent demonstrable practical experience.

Qualifications

Bachelor

Range Of Year Experience-Min Year

15

Range Of Year Experience-Max Year

18

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search