AI Runtime Engineer
Indexed description
AI Runtime Engineer
Vienna, Austria — Hybrid
AI Infrastructure | Distributed Systems | Large Language Models | High-Performance Computing
Our client, an innovative AI Infrastructure company based in Vienna, is building the runtime platform that powers high-performance AI applications used by enterprise customers across Europe.
They're looking for an AI Runtime Engineer to help optimise the systems responsible for serving Large Language Models across distributed GPU infrastructure, where every millisecond of latency and every percentage of GPU utilisation matters.
This is a deeply technical engineering role sitting at the intersection of distributed systems, cloud infrastructure, and modern AI.
Your Responsibilities
- Develop high-performance runtime services responsible for serving production AI models
- Optimise inference performance across distributed GPU infrastructure
- Build scalable model serving systems capable of handling enterprise AI workloads
- Develop platform services and performance tooling using Rust and Python
- Improve scheduling, orchestration, and resource utilisation across Kubernetes clusters
- Work closely with AI Researchers and Machine Learning Engineers to optimise production deployment
- Profile system performance and remove bottlenecks across networking, memory, and compute layers
- Improve platform observability, reliability, and operational efficiency
- Contribute to the architecture of the company's next-generation AI infrastructure platform
Experience Required
- 5+ years of experience in Backend Engineering, Platform Engineering, Distributed Systems, or AI Infrastructure
- Strong commercial experience developing software in Rust or modern systems programming languages
- Excellent Python development skills
- Hands-on experience with Kubernetes in production environments
- Experience deploying or operating large-scale model serving infrastructure
- Good understanding of distributed systems, networking, concurrency, and cloud-native architectures
- Experience working with GPU computing or performance-critical applications
- Passion for solving complex infrastructure and performance engineering challenges
Nice to Have
- Experience with NVIDIA CUDA, Triton Inference Server, or vLLM
- Experience serving Large Language Models in production
- Knowledge of distributed inference frameworks
- Experience with Ray, KServe, or Kubernetes GPU Operators
- Familiarity with observability platforms such as Prometheus, Grafana, or OpenTelemetry
- Previous experience in AI Infrastructure, HPC, or AI Developer Tools companies
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search