Back to search
Confidential Linkedin · Posted 11d ago

Senior Engineer - AI, Inference

Ireland

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

The role

We build the inference layer that powers F5's AI security products: the systems that run large language models fast, cheaply, and reliably at production scale so we can inspect, secure, and govern enterprise giveGenAI traffic in real time.


You'll own how models are served — squeezing maximum throughput out of every GPU, cutting tail latency, and keeping the serving stack observable and self-scaling under real customer load. If you think in tokens-per-second, prefill latency, and GPU memory budgets, this is your seat.


What you'll do

Optimise LLM inference pipelines (LLaMA, GLM, GPT-OSS and similar model classes) for throughput and latency — multi-GPU inference, prefix caching, and memory-efficient serving — targeting order-of-magnitude gains in requests-per-second (RPS).

Own end-to-end model serving in production: deployment, low-latency inference, multi gpu parallelism, and high-throughput serving across cloud and on-prem environments.

Build and maintain OpenAI-compatible serving APIs (e.g. /v1/chat, /v1/responses) that support reliable tool calling across multi-step agentic workflows.

Tune and operate a modern serving stack (vLLM or equivalent) — continuous batching, KV-cache management — balancing throughput, latency, generation quality, and memory footprint.

Maximise GPU across chip architectures like Ada Lovelace, Hopper, Blackwell, etc; profile, benchmark, and eliminate bottlenecks.

Instrument the serving layer with logging, telemetry, and metrics (prefill latency, tokens-per-batch capacity, preemption count, queue depth) to drive observability and metric-based autoscaling.

Ship on Kubernetes: containerised deployments (Docker, Helm), CI/CD, and low-risk, version-controlled rollouts across staging and production.

Benchmark rigorously against recognised standards and build tooling that automates performance characterisation.


What you'll bring (must-have)

• Hands-on production experience serving LLMs at scale, with measurable throughput/latency wins you can walk through end to end.

• Deep familiarity with a modern inference/serving framework (vLLM, TensorRT-LLM, TGI, or similar), including batching and speculative decoding.

• Strong grip on GPU performance: memory management, model/tensor parallelism, and hardware-aware optimisation.

• Solid software engineering in Python OR GoLang, plus Docker + Kubernetes for production deployment.

• A benchmarking mindset — you measure rather than guess, and you can defend the trade-offs.


Nice to have

• Building OpenAI-compatible endpoints and agentic / tool-calling serving paths.

Distributed training exposure (large-model fine-tuning / pretraining on multi-node GPU clusters).

MLPerf submission experience or embedded / ARM inference optimisation.

Release-management / CI-CD ownership across production and staging.

• Interest in AI security — adversarial robustness, model scanning, or securing GenAI in production.


Tech you'll work with

vLLM · NVIDIA H100 / H200 / B200 · · Kubernetes · Docker · Helm · CI/CD · GenAI-Perf

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search