Cloudera Data Engineer
Indexed description
Main Duties And Responsibilities
- Design, build, and maintain big data pipelines using Cloudera CDP components.
- Develop and optimize ETL/ELT workflows using Spark, Hive, and Hadoop ecosystems.
- Manage and maintain HDFS storage, ensuring high availability and performance.
- Perform production support activities, including issue resolution, monitoring, and performance tuning.
- Work with data scientists, analysts, and other engineering teams to support data-driven initiatives.
- Ensure data governance, security, and compliance across big data environments.
- Troubleshoot cluster-level issues related to resource utilization, scheduling, and configuration.
- Participate in solution design, architecture discussions, and technical documentation.
- Strong hands-on experience with Cloudera CDP.
- Proficiency in Hadoop ecosystem: HDFS, YARN, MapReduce, Hive, Impala.
- Strong experience working with Apache Spark (PySpark or Scala preferred).
- Solid understanding of distributed systems and big data architecture.
- Experience with production support and troubleshooting in big data environments.
- Strong SQL skills and experience optimizing large-scale queries.
- Familiarity with Linux, shell scripting, and automation tools.
- Experience with version control (Git) and CI/CD workflows is a plus.
Experience
- Experience with cloud platforms (AWS, Azure, GCP) for big data workloads.
- Exposure to Kafka, NiFi, Oozie, or similar data orchestration tools.
- Knowledge of Docker, Kubernetes, or containerized environments.
- Understanding of security, encryption, and data governance concepts.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search