Computer Vision Engineer
Indexed description
Founding Applied AI Engineer – Agentic Systems & VLMs - Upto £130k + equity
Ever built agents that actually act in the real world, rather than just outputting text in a notebook?
Have you wrestled with Vision-Language Models (VLMs) or Vision-Language-Action models (VLAs) - pushing past raw frame analysis to build systems that observe, reason, decide, and recover when things inevitably go sideways?
Do you thrive on solving messy, real-world problems where bulletproof reliability matters far more than shiny, brittle demos?
If that sounds like you, keep reading.
The Challenge
We’re an early-stage startup building truly autonomous agents that test games the way real players do. No scripted bots. No fragile automation. We’re talking about high-performing agents that observe gameplay, reason over streaming images and video frames, and take continuous action across real devices.
The goal is simple to state, but brutal to execute: make top studios trust autonomous agents more than manual QA.
You’ll be shipping production VLM/VLA-driven agents that operate inside live games across mobile and desktop - navigating inconsistent UIs, timing issues, network lag, and all the edge cases that break naive automation.
This isn't research theatre. This is about getting frontier vision models out of research papers and into production.
What You’ll Be Building
- VLM & VLA Core Systems: Multimodal agents that reason over live image streams and video frames to determine the optimal next action in real time.
- Vision-Driven Autonomy: Agents that interact directly with real games—tapping, swiping, clicking, and navigating dynamically.
- Resilient Agentic Loops: Closed-loop systems that observe, decide, act, and self-heal—strictly avoiding fixed scripts.
- Cross-Platform Scalability: Automation that runs flawlessly across devices, OS versions, screen sizes, and unscripted edge cases.
- Production-Grade Trust: Systems robust enough for studios to run entirely unsupervised.
Think less "prompt engineering", more deep systems architecture.
Why This Role Hits Differently
- Founding-Level Ownership: You’ll directly architect and shape how our core agentic engine is built from day one.
- Real Production Constraints: Zero toy problems. No controlled demos.
- Agents in the Wild: Deploying models that operate autonomously under unpredictable conditions.
- High-Stakes Impact: A product where success lives or dies on pure, unvarnished reliability.
If you get a kick out of seeing your systems break in the wild, diagnosing why, and making them bulletproof - you’ll love it here.
About You
You’re an Applied AI Engineer, Systems Engineer, or Autonomy Specialist who brings:
- Production VLM / VLA Experience: Direct experience building and deploying multimodal or vision-based agents that take actions using images and video (not just text).
- Strong Computer Vision Foundations: Familiarity with classic and modern CV architectures (YOLO, ResNet, EfficientNet, etc.).
- Frontier Model Familiarity: Solid understanding of Vision Agents and VLMs (LLaVA, CLIP, Flamingo, or custom VLA pipelines).
- Messy Domain Exposure: Experience in high-complexity technical environments—robotics, device control, autonomous vehicles, complex UI automation, or real-time control.
- User-Centric Collaboration: Comfort working alongside QA Leads and Testers to deeply understand and automate their real workflows.
- 0-to-1 Startup Grit: Experience working in fast-moving AI product startups, ideally building early-stage infrastructure from scratch.
Ready to Build?
If you’ve taken multimodal models into production and want real founding ownership over a ground-breaking product, let’s talk.
Apply directly or get in touch:
- ✉️ Email: [email protected]
- 📞 Phone: 020 3773 6498
Note: You must be eligible to work in the UK or EU.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search