Back to search
UST Linkedin · Posted 5d ago

6 + YoE - Data Engineer – Big Data / PySpark - any UST Location - Immediate Joiner

India

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Candidates ready to join immediately can share their details via email for quick processing.

📌 CCTC | ECTC | Notice Period | Location Preference

[email protected]

Act fast for immediate attention! ⏳📩

___________________________________________________________________________________________________

Must-Have Skills

  • 6+ years of overall experience in Data Engineering / Big Data.
  • Strong understanding of Big Data concepts and architecture.
  • Strong hands-on experience with Apache Spark.
  • Expertise in:
  • Spark Performance Tuning
  • Spark Optimization
  • Query/Job Performance Improvement
  • Troubleshooting Spark workloads
  • Strong hands-on experience with PySpark and Spark.
  • Strong programming experience in Python.
  • Good experience working with MySQL / SQL.
  • Strong experience in designing and developing Data Pipelines.
  • Hands-on experience with Apache Airflow for data pipeline orchestration and scheduling.
  • Experience working with at least one Cloud Platform.
  • GCP experience is preferred.
  • Good understanding of CI/CD and DevOps concepts.
  • Experience integrating data engineering workloads with CI/CD pipelines.
  • Strong debugging, troubleshooting, and problem-solving skills.

Preferred Skills

  • Hands-on exposure to relevant GCP data services.
  • Experience handling large-scale and high-volume datasets.
  • Understanding of distributed data processing and data architecture.
  • Experience improving the scalability, reliability, and performance of data pipelines.
  • Exposure to Agile development and DevOps practices.

Key Responsibilities

  • Design, develop, and maintain scalable Big Data and Data Engineering solutions.
  • Develop data processing applications using Python, PySpark, and Apache Spark.
  • Perform Spark performance tuning and optimization for large-scale workloads.
  • Build, maintain, and monitor robust ETL/ELT data pipelines.
  • Develop and manage workflow orchestration using Apache Airflow.
  • Work with MySQL/SQL for data extraction, transformation, and validation.
  • Deploy and support data engineering solutions in cloud environments, preferably GCP.
  • Work with DevOps teams to implement and maintain CI/CD pipelines.
  • Troubleshoot production issues and optimize data processing performance.
  • Collaborate with engineering and business teams to deliver reliable and scalable data solutions.

Primary Skill Combination

Big Data + Apache Spark + PySpark + Python + Airflow + SQL/MySQL + Cloud (GCP Preferred) + CI/CD

Mandatory Focus: Strong hands-on Apache Spark performance tuning and optimization experience.


Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search