Back to search
XpertDirect Linkedin · Posted 4d ago

AI Runtime Engineer

Vienna

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

AI Runtime Engineer

Vienna, Austria — Hybrid

AI Infrastructure | Distributed Systems | Large Language Models | High-Performance Computing


Our client, an innovative AI Infrastructure company based in Vienna, is building the runtime platform that powers high-performance AI applications used by enterprise customers across Europe.

They're looking for an AI Runtime Engineer to help optimise the systems responsible for serving Large Language Models across distributed GPU infrastructure, where every millisecond of latency and every percentage of GPU utilisation matters.

This is a deeply technical engineering role sitting at the intersection of distributed systems, cloud infrastructure, and modern AI.


Your Responsibilities

  • Develop high-performance runtime services responsible for serving production AI models
  • Optimise inference performance across distributed GPU infrastructure
  • Build scalable model serving systems capable of handling enterprise AI workloads
  • Develop platform services and performance tooling using Rust and Python
  • Improve scheduling, orchestration, and resource utilisation across Kubernetes clusters
  • Work closely with AI Researchers and Machine Learning Engineers to optimise production deployment
  • Profile system performance and remove bottlenecks across networking, memory, and compute layers
  • Improve platform observability, reliability, and operational efficiency
  • Contribute to the architecture of the company's next-generation AI infrastructure platform


Experience Required

  • 5+ years of experience in Backend Engineering, Platform Engineering, Distributed Systems, or AI Infrastructure
  • Strong commercial experience developing software in Rust or modern systems programming languages
  • Excellent Python development skills
  • Hands-on experience with Kubernetes in production environments
  • Experience deploying or operating large-scale model serving infrastructure
  • Good understanding of distributed systems, networking, concurrency, and cloud-native architectures
  • Experience working with GPU computing or performance-critical applications
  • Passion for solving complex infrastructure and performance engineering challenges


Nice to Have

  • Experience with NVIDIA CUDA, Triton Inference Server, or vLLM
  • Experience serving Large Language Models in production
  • Knowledge of distributed inference frameworks
  • Experience with Ray, KServe, or Kubernetes GPU Operators
  • Familiarity with observability platforms such as Prometheus, Grafana, or OpenTelemetry
  • Previous experience in AI Infrastructure, HPC, or AI Developer Tools companies
Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search