Machine Learning / Algorithm Engineer
Indexed description
Role Direction:
This role requires processing large-scale data — including high-frequency tick data, user behavior data, trading data, fund data, and more — to provide stable, accurate, and reusable data and feature support for risk control models, anomaly detection, user risk scoring, and real-time monitoring.
Key Responsibilities:
- Build risk control data pipelines — covering data ingestion, cleaning, processing, validation, storage, and output
- Process large-scale financial data — including high-frequency tick data, market data, trading behavior data, user behavior data, fund transaction records, and more
- Establish a risk control feature engineering system — providing high-quality features for ML models, risk scoring models, and anomaly detection models
- Support model training and deployment — providing stable datasets, feature tables, labeling systems, and data validation logic
- Collaborate with Data Scientists, Risk Control teams, and Backend Engineers to integrate risk control data capabilities into production business systems
- Optimize large-scale data processing performance — improving efficiency, accuracy, and stability
- Contribute to risk control data monitoring systems — including data quality monitoring, anomaly fluctuation detection, and feature stability monitoring
Required Qualifications:
- 2+ years of experience in financial data, risk control data, data engineering, or large-scale user data processing
- Proficient in Python and SQL — able to independently handle complex data processing, cleansing, feature construction, and performance optimization
- Familiar with data processing and feature engineering methods — able to transform raw business data into structured features for risk control models
- Experience with high-frequency tick data, market data, trading data, or financial time-series data processing
- Experience with user behavior data processing — understanding user journeys, behavior logs, event tracking, conversion funnels, and related data structures
- Basic understanding of machine learning model principles — able to understand model requirements for data, features, labels, and sample quality
- Strong data quality awareness — able to identify and handle issues such as missing values, duplicates, anomalies, delays, and inconsistent definitions
- Capable of handling high-concurrency, high-frequency, large-volume data — with attention to pipeline stability and scalability
Preferred Qualifications:
- Experience in FinTech, trading platforms, brokerages, payments, anti-fraud, AML, or risk control systems
- Familiarity with financial business data — including deposits, withdrawals, trading, positions, accounts, KYC, user behavior, market data, and more
- Experience with real-time data processing — familiarity with Kafka, Flink, Spark, Airflow, DBT, or similar tools is a plus
- Experience with Feature Store, data warehouses, metric systems, labeling systems, or risk control data platform development
- Familiarity with model feature monitoring, data drift detection, sample construction, feature backfilling, training dataset generation, and related workflows
- Experience with large-scale customer data processing — able to support data analysis and modeling for millions or more users
What We're Looking For:
- Not just extracting data and writing SQL — able to understand the business logic behind risk control scenarios
- High standards for data accuracy, stability, and traceability
- Able to organize complex, distributed, real-time financial data into modelable, monitorable, and reusable data assets
- Able to collaborate effectively with Algorithm, Risk Control, Product, and Engineering teams — driving data capabilities into production business systems
岗位定位:
该岗位需要处理高频 Tick 数据、用户行为数据、交易数据、资金数据等大规模数据,并为风控模型、异常检测、用户风险评分和实时监控提供稳定、准确、可复用的数据与特征支持。
岗位职责:
- 负责风控相关数据链路建设,包括数据采集、清洗、加工、校验、存储和输出
- 处理高频 Tick 数据、行情数据、交易行为数据、用户行为数据、资金流水等大规模金融数据
- 建立风控特征工程体系,为机器学习模型、风险评分模型、异常检测模型提供高质量特征
- 支持风控模型训练和上线,提供稳定的数据集、特征表、标签体系和数据验证逻辑
- 与数据科学家、风控团队、后端工程师协作,将风控数据能力接入实际业务系统
- 优化大规模数据处理性能,提升数据处理效率、准确性和稳定性
- 参与风控数据监控体系建设,包括数据质量监控、异常波动检测、特征稳定性监控等
任职要求:
- 2 年以上金融数据、风控数据、数据工程或大规模用户数据处理经验
- 精通 Python 和 SQL,能够独立完成复杂数据处理、数据清洗、特征构建和性能优化
- 熟悉数据处理和特征工程方法,能够将原始业务数据转化为可用于风控模型的结构化特征
- 有高频 Tick 数据、行情数据、交易数据或金融时间序列数据处理经验
- 有用户行为数据处理经验,理解用户路径、行为日志、事件埋点、转化漏斗等数据结构
- 了解机器学习模型基本原理,能理解模型对数据、特征、标签和样本质量的要求
- 具备较强的数据质量意识,能发现并处理缺失、重复、异常、延迟、口径不一致等问题
- 能够处理高并发、高频率、大体量数据,并关注数据链路的稳定性和可扩展性
加分项:
- 有金融科技、交易平台、券商、支付、反欺诈、反洗钱、风控系统相关经验
- 熟悉入金、出金、交易、持仓、账户、KYC、用户行为、行情等金融业务数据
- 有实时数据处理经验,熟悉 Kafka、Flink、Spark、Airflow、DBT 等工具优先
- 有 Feature Store、数据仓库、指标体系、标签体系或风控数据平台建设经验
- 熟悉模型特征监控、数据漂移、样本构建、特征回溯、训练集生成等流程
- 有大规模客户数据处理经验,能够支持百万级或更高体量用户数据分析与建模
我们希望你是这样的人:
- 不只是取数和写 SQL,而是能理解风控场景背后的业务逻辑
- 对数据准确性、稳定性和可追溯性有高要求
- 能把复杂、分散、实时的金融数据整理成可建模、可监控、可复用的数据资产
- 能与算法、风控、产品和研发团队高效协作,推动数据能力真正落地到业务系统中
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search