Consultant Specialist
Indexed description
We are currently seeking an experienced professional to join our team in the role of Consultant Specialist.
Business: CTO Data Technology
Principal Responsibilities
- Design and deliver CDMS migration and incremental sync solutions: initial full load, CDC cutover, replay/backfill and recovery mechanisms
- Develop and maintain batch/stream processing jobs using PySpark / Spark SQL / Structured Streaming (cleansing, standardization, mapping, deduplication and merge)
- Integrate with Kafka CDC events: topic/partition/key strategy, handle out-of-order/duplicate/late events, replay, and idempotent processing
- Migrate and incrementally ingest MongoDB data: document parsing, schema evolution, nested/array flattening, and mapping to target models
- Collaborate with Java / Spring Boot service teams on event contracts (JSON/Avro), schema management, error handling, retry and DLQ patterns
- Implement data landing and modelling on the Hadoop ecosystem (HDFS/Hive and related storage formats), including partitioning and small-file management
- Orchestrate pipelines with Airflow: DAG design, dependency management, parameterisation, SLA/alerting, rerun and backfill procedures
- Build data quality and reconciliation controls: field-level validation, key uniqueness, hash/sample checks, reconciliation reports and audit evidence
- Optimise performance and stability: shuffle/skew handling, state management & checkpointing, resource tuning, backpressure and fault tolerance
- Contribute to engineering excellence: code reviews, standards, documentation, production support and continuous improvement
- Strong Python fundamentals and engineering practices; solid hands-on PySpark experience (DataFrames, Spark SQL, RDD, tuning and troubleshooting)
- Proven experience with Spark Structured Streaming: watermarking, stateful processing, checkpointing, and delivery semantics (at-least-once / exactly-once concepts)
- Strong Kafka experience: topic design, partitioning/keying, consumer groups, offset management, replay, ordering and idempotency
- Hands-on CDC implementation experience: initial load, incremental continuity, backfill, restorability and consistency guarantees
- Experience operating on on prem Hadoop (HDFS/Hive/YARN) and resolving common runtime issues
- Airflow experience: DAG patterns, parameterization, dependency/SLA management, alerting, retries and backfills
- Working knowledge of MongoDB: data modelling, querying/aggregation, indexing; understanding of change streams/oplog concepts
- Working knowledge of Java (able to read code, troubleshoot, and align on interfaces; Spring Boot is a plus)
- Strong SQL skills and practical experience in data reconciliation and data quality controls
- Ability to understand requirements in written English and translate them into implementable technical designs (HLD/LLD)
Personal data held by the Bank relating to employment applications will be used in accordance with our Privacy Statement, which is available on our website.
***Issued By HSBC Software Development (GuangDong) Limited***
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search