Computer Vision Engineer
Indexed description
We work with marquee customers including Tata Steel, JSW, ArcelorMittal, Vedanta, Godrej & Boyce, Grasim, Holcim, and Jindal Steel, and are scaling globally across India, the Middle East, Europe, and North America. We are one of the few Indian AI product start-ups to be a partner to GCP, Azure, and AWS, and the AI partner of choice for CII, ICC, and NASSCOM.
The Role
We are looking for a hands-on Computer Vision Engineer to work on the models that power Ripik's industrial AI platform. You will own CV problems end-to-end — from data strategy and annotation to model development, edge deployment, and production monitoring — for some of the most complex vision problems in heavy industry.
Key Responsibilities
- Own computer vision problems end-to-end — from problem framing and data strategy through model development, edge deployment, and production monitoring — across Ripik's industrial portfolio (steel, cement, pharma, paints, and beyond).
- Build models for hard vision challenges — novel defect types, extreme class imbalance, multi-camera fusion, low-light / high-noise factory environments, and real-time inference on constrained edge hardware.
- Stay at the cutting edge of CV research and rapidly evaluate and adopt new models and techniques — YOLO26, SAM 3, Vision Transformers (DINOv2, Swin), Grounding DINO, RF-DETR, zero-shot / open-vocabulary detection (YOLO-World, CLIP) — translating papers into production value.
- Follow and contribute to engineering standards for the vision stack — model training pipelines, data versioning (DVC), annotation workflows (CVAT, Roboflow, Label Studio), experiment tracking (W&B, MLflow), edge export formats (TensorRT, ONNX, OpenVINO), and CI/CD for model updates.
- Drive inference optimisation — quantisation (INT8 / FP16, GPTQ), pruning, knowledge distillation, and batching strategies — to meet latency and cost targets across NVIDIA Jetson, industrial PCs, and cloud GPU instances.
- Debug production issues on live customer deployments — trace performance drops, root-cause failure modes, and ship fixes with the right guardrails.
- Champion a data-centric AI approach — invest in annotation quality, active learning, synthetic data generation, and feedback loops from production rather than only chasing bigger models.
- Build robust evaluation frameworks — domain-specific metrics, A/B testing against production baselines, and systematic failure-mode analysis to ensure models deliver real business impact.
- Partner cross-functionally with product, field engineering, operations, and leadership — translate business problems into well-scoped modelling projects and communicate results clearly.
- Bachelor's in Computer Science, AI/ML, Electrical Engineering, or a related field.
- 1–3 years of hands-on experience in computer vision — with a strong track record of taking models from research / prototyping through to production deployment.
- Deep proficiency in Python and PyTorch; strong working knowledge of OpenCV, Albumentations, and image / video processing fundamentals.
- Demonstrated expertise across multiple CV tasks — object detection, instance / semantic / panoptic segmentation, anomaly detection, pose estimation, or tracking.
- Hands-on experience with modern model families — YOLO (v8 / v11 / v26), transformer-based detectors (RT-DETR, DETR, RF-DETR), segmentation models (SAM / SAM 2), and CNN backbones (ResNet, EfficientNet, ConvNeXt, Vision Transformers).
- Production experience deploying models to edge or on-prem hardware using TensorRT, ONNX Runtime, or OpenVINO; comfort with Docker, Kubernetes, and at least one cloud platform (AWS / Azure / GCP).
- Strong first-principles problem-solving — comfortable navigating novel, unstructured problems where no playbook exists.
- Experience in a high-growth start-up or similarly fast-paced environment.
- Excellent communication — able to distil complex technical concepts for non-technical stakeholders, write clear documentation, and present results to leadership and customers.
- Experience with industrial or manufacturing domains — understanding of factory-floor constraints, camera setups, lighting variability, and integration with PLCs / SCADA systems.
- Familiarity with zero-shot and open-vocabulary detection (Grounding DINO, YOLO-World, CLIP) and foundation models (DINOv2, SAM 3, Florence) for data-efficient learning.
- Exposure to vision–language models (GPT-4o vision, Gemini, LLaVA) for combining visual inspection with natural-language reporting or operator copilots.
- Knowledge of 3D vision, depth estimation, point-cloud processing, or multi-camera calibration for volumetric industrial inspection.
- Experience with multi-object tracking (ByteTrack, BoT-SORT) and video analytics pipelines for continuous production-line monitoring.
- Contributions to open-source CV projects, publications in top-tier venues (CVPR, ECCV, ICCV, NeurIPS), or strong Kaggle competition results.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search