Principal AI Engineer – Computer Vision & Video AI
Indexed description
📍Al Khobar, Saudi Arabia
→ Up to SAR 50k per month
We’re partnering with a technology company using AI to improve safety across large-scale industrial and construction environments. Their platform analyses live site video to identify potential hazards and provide teams with timely, actionable safety insights.
They’re looking for a Principal AI Engineer to take technical ownership of the AI and perception stack. You’ll stay hands-on while setting architecture, solving the hardest computer vision and applied AI problems, and influencing technical direction across the wider engineering team.
This is a genuinely deep technical role where video is at the centre of the product, spanning computer vision, multimodal models, agentic AI, edge deployment and real-time systems.
What You’ll Do:
- Own the technical direction of the computer vision and perception stack.
- Design and build production systems for object detection, segmentation, tracking and video understanding.
- Develop multimodal and vision-language systems that can reason over live and recorded video.
- Build AI agents capable of reasoning over visual data, using tools and APIs and operating within controlled workflows.
- Own the full model lifecycle from problem definition and data strategy through to experimentation, deployment, monitoring and continuous improvement.
- Optimise models across edge and cloud environments, balancing accuracy, latency, reliability and cost.
- Establish responsible AI and safety controls, including confidence thresholds, human review, explainability and auditability.
- Lead architecture reviews and document key technical decisions.
Who We’re Looking For:
- Proven experience taking multiple production ML or computer vision systems from concept through to live operation.
- Deep hands-on experience working with video, including video streaming, analytics or computer vision over video.
- Strong experience with PyTorch or TensorFlow across detection, segmentation, tracking and video understanding.
- Experience with both purpose-built computer vision models and transformer-based approaches.
- Production experience with vision-language or multimodal models, including model adaptation, prompting, evaluation and integration.
- Experience building agentic AI or tool-using systems, with appropriate guardrails and human-in-the-loop controls.
- Strong experience with dataset curation, fine-tuning, active learning and experiment management.
- Proven experience deploying AI workloads across edge and cloud infrastructure, with a focus on real-world latency, throughput and memory constraints.
- Strong MLOps and evaluation experience, including regression testing, experiment tracking, model versioning, monitoring and drift detection.
- Exceptional written and spoken English.
Nice to have:
- RAG, vector search or multimodal retrieval experience.
- Real-time video or streaming camera systems.
- Edge AI technologies such as TensorRT, ONNX Runtime, Jetson or other GPU/NPU platforms.
- Self-hosted LLM/VLM serving using tools such as vLLM or similar.
- Experience in safety-critical, industrial or construction environments.
- Knowledge of EHS or workplace safety frameworks.
- Experience with privacy-focused video systems, including anonymisation, access controls and audit trails.
Interested?
Apply today!
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search