Kafka / Spark Software Development Lead
Indexed description
Kafka / Spark Software Development Lead
Location: Pittsburgh, PA / Cleveland, OH
Work Arrangement: Onsite – 5 days per week
Experience: 8+ years
Employment Type: W2 / Direct Hire
Work Authorization: U.S. Citizens, Green Card holders, and candidates eligible for H-1B transfer
Sponsorship: Not available
About the Role
We are seeking an experienced Kafka / Spark Software Development Lead to join our Applications Development and Maintenance team supporting a large U.S. banking client.
The ideal candidate will have strong hands-on experience designing and developing large-scale, real-time data streaming and distributed processing solutions using Apache Kafka and Apache Spark. This role requires strong technical leadership, software development, system design, troubleshooting, and stakeholder management skills.
Key Responsibilities
- Design, develop, and implement scalable real-time data streaming solutions using Apache Kafka and Spark Structured Streaming.
- Build and enhance Kafka producers, consumers, topics, partitions, and event-driven architectures for reliable and high-throughput data ingestion.
- Develop and optimize Spark Streaming / Spark Structured Streaming applications for real-time transformation, aggregation, enrichment, and analytics.
- Integrate Kafka and Spark with data lakes, data warehouses, relational databases, NoSQL databases, APIs, cloud platforms, and enterprise applications.
- Implement highly available and resilient streaming pipelines using checkpointing, replication, schema management, fault tolerance, and recovery mechanisms.
- Monitor, troubleshoot, and tune Kafka and Spark streaming applications to improve performance, scalability, throughput, and reliability.
- Analyze business and technical requirements and translate them into scalable technical solutions and system designs.
- Create technical designs for new systems and enhancements to existing applications.
- Define technical scope, assumptions, dependencies, and implementation approaches for assigned initiatives.
- Collaborate with solution architects, enterprise architects, data engineers, DevOps teams, product owners, and business stakeholders.
- Provide technical leadership and guidance to development teams on streaming architecture and distributed data processing.
- Participate in Agile/Scrum ceremonies and contribute to estimation, prioritization, planning, and delivery.
- Establish and maintain effective working relationships with clients, project teams, and cross-functional stakeholders.
- Research emerging technologies and identify opportunities to improve enterprise data streaming and event-driven platforms.
- Ensure solutions adhere to enterprise architecture, security, coding, and development standards.
Required Qualifications
- 8+ years of experience designing, developing, and supporting large-scale distributed data processing and streaming applications.
- Strong hands-on experience with Apache Kafka, including:
- Kafka Topics and Partitions
- Producers and Consumers
- Consumer Groups
- Kafka Connect
- Schema Registry
- Replication and Fault Tolerance
- Extensive experience with Apache Spark Streaming and/or Spark Structured Streaming.
- Strong programming experience in at least one of the following:
- Java
- Scala
- Python / PySpark
- Strong understanding of distributed systems, message-oriented middleware, data partitioning, fault tolerance, scalability, and high availability.
- Experience integrating Kafka and Spark with RDBMS, NoSQL databases, data lakes, data warehouses, cloud storage, APIs, and enterprise applications.
- Strong experience in performance tuning, troubleshooting, monitoring, and optimization of distributed streaming applications.
- Experience with event-driven architecture and real-time data processing.
- Strong understanding of software development lifecycle, coding standards, testing, deployment, and production support.
- Experience working in Agile/Scrum development environments.
- Strong technical leadership, analytical, problem-solving, communication, and stakeholder management skills.
Preferred Qualifications
- Experience working on large-scale banking or financial services platforms.
- Experience with cloud-based data platforms such as AWS, Azure, or GCP.
- Experience with Kafka Streams, Confluent Platform, or similar streaming technologies.
- Experience with CI/CD, DevOps, containerization, and automated deployment pipelines.
- Experience working with enterprise-scale data architecture and modernization initiatives.
Interested candidates can share their profile at [email protected]
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search