Back to search
HSBC SOFTWARE DEV (GD) LTD Linkedin · Posted 3d ago

AI Data Engineer

Shenzhen

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Principal responsibilities:

• Build and operate pipelines for structured domain data (onboarding events, customer/entity data, credit facilities/exposures, limits, risk grades, portfolio hierarchies, performance and arrears).

• Build unstructured data pipelines for domain documents (KYC documents, corporate registries extracts, credit memos, financial statements, portfolio review materials): parsing, metadata enrichment, deduplication and retention handling.

• Develop embedding/vectorisation pipelines and manage vector indices with refresh and deletion strategies aligned to data governance.

• Implement data quality, observability and lineage: automated testing, SLAs, anomaly detection, monitoring and runbooks.

• Apply privacy-by-design: PII handling, masking/tokenisation where required, access controls and audit trails.

• mplement dataset versioning and reproducibility (snapshots, schema evolution, data contracts) to support repeatable evaluations.

• Partner with Agent Engineers and stakeholders to define data requirements and evaluation datasets; close feedback loops from production retrieval/agent performance.

• Support production operations and continuous improvement of data services powering AI.


Knowledge & Experience / Qualifications:

• Delivering reliable, fresh, well-governed datasets and indices used by KYC, credit and portfolio AI solutions.

• Improving retrieval relevance and answer grounding by strengthening metadata, chunking inputs, and data quality.

• Reducing operational incidents through strong testing, monitoring, and disciplined change management.

• Making AI datasets reproducible and auditable to support reviews and control expectations.


What additional skills will be good to have? :

• Experience working with sensitive data domains (KYC/AML, credit risk) and implementing controlled access patterns.

• Familiarity with vector search concepts, document ranking/reranking and retrieval metrics.

• Experience with feature stores or MLOps-style dataset management.

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search