Machine Learning Engineer
Indexed description
You will implement and productionize modern CV/ML pipelines across two tracks: (1) video understanding with personalization and synthetic data generation, and (2) real-time multi-camera detection/tracking with dataset auto-labeling. The focus is practical – training and evaluation, data quality loops, and inference performance.You take a module from prototype to stable pipeline with clear metrics.
Key ProjectsTrack A – Video understanding, keypoints, personalization, synthetic data
Core stack centers on video transformers and/or 3D CNN baselines plus pose/hand keypoint pipelines (TimeSformer, MViT, VideoMAE; MediaPipe Hands; MMPose/RTMPose; ViTPose).
Personalization is implemented via parameter-efficient adaptation (e.g., LoRA/adapters) and signer-specific calibration layers on top of a shared backbone.
Synthetic data pipeline leverages diffusion-based video generation frameworks (Stable Video Diffusion, AnimateDiff, VideoCrafter2, CogVideoX) to expand edge cases and improve generalization.
Track B – Real-time detection/tracking, optimization, auto-labeling
Primary detection families target real-time accuracy/latency tradeoffs (YOLOv10; RT-DETR.
Tracking and temporal association uses modern MOT baselines (ByteTrack, BoT-SORT).
Auto-labeling/bootstrapping relies on open-vocabulary detection + promptable segmentation pipelines.
Responsibilities- Build training/inference pipelines for video understanding using transformer backbones (TimeSformer, MViT, VideoMAE) and/or 3D CNN baselines as needed.
- Implement keypoint-driven features and models using Google MediaPipe Hands and OpenMMLab MMPose (RTMPose, ViTPose) to improve robustness under viewpoint/lighting variation.
- Implement personalization layers (LoRA/adapters, calibration heads) and evaluation protocols for user-specific performance without catastrophic forgetting.
- Design and run synthetic data generation loops using diffusion video models with strict QA gates (artifact detection, distribution checks).
- Train, evaluate, and optimize real-time detectors (YOLO, RT-DETR), including ablations and latency profiling.
- Implement multi-object tracking pipelines and tune association logic for occlusions and crowded scenes.
- Build auto-labeling workflows using Grounding DINO + SAM, including human-in-the-loop review and active learning sampling.
- Ship reproducible experiments (configs, seeds, dataset versions), write tests for data/model logic, and document failure modes and monitoring signals.
- 3-5 years of practical ML/CV engineering experience with at least one production or production-like pipeline shipped.
- Strong Python + PyTorch (preferred) or TensorFlow; ability to write and debug custom training/eval loops.
- Hands-on experience with video modeling families (at least one of TimeSformer/MViT/VideoMAE-style pipelines).
- Experience with real-time detection architectures (YOLO and/or RT-DETR class systems).
- Familiarity with pose/hand keypoints tooling (MediaPipe Hands and/or MMPose/RTMPose/ViTPose).
- Solid understanding of training mechanics: augmentation, optimization, mixed precision (FP16), metric-driven iteration.
- Inference and deployment fundamentals: exporting (ONNX), GPU acceleration (TensorRT/CUDA/OpenVINO) and profiling.
- Comfort with Linux + Docker-based workflows.
- Solid data engineering hygiene: dataset QC, train/val split discipline, reproducibility, and experiment tracking.
- Experience with open-vocabulary detection + segmentation for labeling/bootstrapping (Grounding DINO , SAM).
- Experience with video diffusion models for synthetic data generation (Stable Video Diffusion, AnimateDiff, VideoCrafter2, CogVideoX).
- OCR/label-reading components when needed (e.g., PaddleOCR).
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search