Back to search
Discovered MENA Linkedin · Posted 2d ago

Principal AI Ops Engineer

United Arab Emirates

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Principal AI Ops Engineer – Abu Dhabi

Discover the Opportunity:

We’re partnering with a leading organisation in Abu Dhabi that is building and operating advanced AI systems at significant scale.


They’re looking for a Principal AI Ops Engineer to take ownership of how AI and LLM systems operate in production, covering inference and model serving, deployment, observability, reliability and performance.


This is a Principal-level individual contributor role for someone who combines deep AI infrastructure expertise with strong software and reliability engineering fundamentals. You’ll set the operational standards that allow engineering teams to deploy and run production AI systems safely, reliably and efficiently.


Discover the Responsibilities:

  • Design, operate and optimise GPU-based inference and model-serving infrastructure for production AI and LLM workloads.
  • Optimise model serving across latency, throughput, batching, quantisation, autoscaling and infrastructure cost.
  • Build automated release pipelines for models, prompts and agent configurations, including canary deployments, regression gates and rollback strategies.
  • Establish AI-specific observability across model, retrieval and orchestration layers, including tracing, latency, cost and quality monitoring.
  • Define and maintain SLOs across availability, latency and AI system quality, alongside automated alerting and incident response processes.
  • Own capacity planning and cost optimisation across GPU and AI infrastructure.
  • Build secure, scalable Kubernetes environments and Infrastructure-as-Code patterns for production AI workloads.
  • Develop reusable deployment patterns, tooling and operational standards that enable engineering teams to ship AI systems reliably.
  • Lead complex production incidents, load testing and root-cause analysis across AI infrastructure and applications.
  • Provide technical leadership and help establish engineering standards for operating AI systems at scale.


Discover the Requirements:

  • Proven experience operating at Staff, Principal or equivalent senior IC level, with a track record of running production ML or LLM systems at scale.
  • Deep hands-on experience with GPU-based inference and model serving, including technologies such as vLLM, TGI, TensorRT-LLM or similar.
  • Strong understanding of batching, quantisation, latency, throughput, autoscaling and the performance trade-offs involved in production LLM serving.
  • Strong experience with AI/LLM observability, including tracing, quality monitoring, drift and regression detection.
  • Strong reliability engineering fundamentals across SLOs, incident response, capacity planning and post-mortems.
  • Strong Python engineering skills with experience building production-grade automation and infrastructure tooling.
  • Deep experience with Kubernetes, Docker and Infrastructure as Code, ideally Terraform, across cloud environments.
  • Experience with observability technologies such as Langfuse, LangSmith, Arize Phoenix, Grafana or Prometheus.
  • Experience integrating AI evaluation and regression testing into CI/CD and production release processes.
  • Experience with cloud infrastructure, ideally Azure, and production environments with strong security, data residency or compliance requirements.
  • Experience with GPU/AI infrastructure cost optimisation and FinOps would be advantageous.
  • A highly hands-on approach, with the ability to set technical standards while remaining close to engineering and production systems.
Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search