Research Engineer / Data Scientist: Data-Centric AI & Model Evaluation
Indexed description
Research Engineer / Data Scientist: Data-Centric AI & Model Evaluation
Seoul Gangnam (On-site)
- Full-time·Talent Pool / Year-round Interest
You will define task taxonomies, benchmarks, metrics, annotation strategies, and experiment designs across perception, robot learning, VLA, world models, and end-to-end autonomy. You will connect offline model results, simulation outcomes, and real-robot runs into a decision system for research priorities and release quality.
Key Areas
Evaluation designFailure taxonomyData qualityExperimentationBenchmarkingError analysisRelease criteria
What You'll Do
- Define task, scenario, and failure taxonomies with metrics that reflect safety, reliability, efficiency, and recovery.
- Build evaluation datasets and harnesses spanning offline inference, simulation, replay, hardware-in-the-loop, and real-robot runs.
- Analyze slice-level performance, causal drivers, regressions, calibration, and statistical uncertainty.
- Design annotation guidelines, sampling strategies, active-learning loops, and data-quality audits.
- Create scorecards and release gates that help researchers and engineering leaders choose the next experiment and decide when a model is ready.
- BS, MS, or Ph.D. in Machine Learning, Statistics, Data Science, Robotics, Computer Science, or equivalent experience.
- Strong foundation in statistics, experimental design, machine learning evaluation, and data analysis.
- Strong Python and SQL skills and experience building reproducible analyses or evaluation systems.
- Ability to translate ambiguous behavior into measurable definitions without losing operational meaning.
- Excellent communication skills for explaining uncertainty, tradeoffs, and failure patterns to research and engineering teams.
- Experience evaluating robotic manipulation, autonomous systems, multimodal models, or safety-critical ML.
- Experience with causal inference, sequential experiments, uncertainty calibration, or human evaluation.
- Experience building annotation operations, data-quality programs, or active-learning systems.
- Experience with dashboards, experiment registries, model cards, or release governance.
- Experience linking offline metrics to field outcomes and identifying misleading proxies.
- Competitive compensation based on experience, level, and technical impact.
- The tools, compute, robots, and equipment needed to do your best work.
- The opportunity to build, test, and deploy directly on FRIDAY and the Holiday Robotics full stack.
- Close collaboration with researchers and engineers across hardware, control, simulation, learning, perception, data, and operations.
Apply
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search