Java Spark Engineer – Berkeley Heights, NJ (Full‑time, Onsite)
Indexed description
Job Title: Java Spark Engineer
Location: Berkeley Heights, NJ (Onsite, 5 days/week – flexible to support weekends)
Type: Full‑time (Permanent) Experience Level: Senior (10+ years)
Job Description:
Role Overview
We are seeking a Java Spark Engineer to architect and build scalable, fault‑tolerant data pipelines. The candidate will lead design of batch and streaming ETL/ELT systems, optimize performance, and ensure production reliability. This role requires strong expertise in Java, Apache Spark, distributed systems, and cloud‑managed clusters, along with leadership and mentoring skills.
Key Responsibilities
- Architect and build scalable, fault‑tolerant data pipelines using Apache Spark (Java)
- Lead design of batch and streaming ETL/ELT systems handling large data volumes
- Perform deep‑dive performance tuning (partitioning, memory management, shuffle/skew optimization, job cost reduction)
- Set coding standards and lead code/design reviews across the team
- Drive technical decisions on data architecture, storage formats, and pipeline orchestration
- Mentor mid‑level and junior engineers; act as technical escalation point
- Partner with product, analytics, and platform teams to translate requirements into scalable systems
- Own production reliability (on‑call ownership, incident response, root‑cause analysis)
- Evaluate and introduce new tools/frameworks to improve systems
- Contribute to capacity planning and cost optimization for cluster infrastructure
Required Qualifications
- Bachelor’s/Master’s in Computer Science, Engineering, or related field
- 7+ years professional Java development experience
- 5+ years hands‑on Apache Spark in production environments
- Expert understanding of distributed systems (fault tolerance, data locality, shuffle mechanics, resource management)
- Proven track record designing systems processing terabyte+ scale data
- Strong SQL skills and familiarity with columnar storage formats (Parquet, ORC, Avro, Delta Lake/Iceberg)
- Experience with cluster managers (YARN, Kubernetes) and cloud‑managed Spark
- Proficiency with Kafka
- Strong grasp of CI/CD, containerization, and infrastructure‑as‑code practices
Preferred Qualifications
- Experience with Flink or other stream‑processing frameworks
- Familiarity with data governance, lineage, and quality frameworks
- Workflow orchestration at scale (Airflow, Oozie, etc.)
- System design for multi‑tenant/multi‑region data platforms
- Prior experience leading a team or acting as a technical lead
Soft Skills / Leadership
- Excellent communication — able to explain technical tradeoffs to non‑technical stakeholders
- Strong mentorship and coaching ability
- Comfortable driving ambiguous, cross‑team technical initiatives
- Team player with strong problem‑solving and analytical skills
- Ability to work independently or within a team
- Adaptable to changing environments
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search