Back to search
YTL AI Labs Linkedin · Posted 23d ago

Model Performance Engineer

Kuala Lumpur

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

About YTL AI Labs

At YTL AI Labs, we build sovereign AI models that perform on par with the world’s best—while staying grounded in local needs, values, and context. Our flagship model, Ilmu, is designed to be culturally aware, contextually intelligent, and fluent in Bahasa Melayu, delivering cutting-edge solutions that empower Malaysian businesses with intelligence that truly understands the market and the people they serve.

As pioneers of sovereign AI, we believe every nation should have the power to shape its own intelligence—guided by its people, priorities, and principles.


Role Overview

We are seeking a Model Performance Engineer to lead the design and optimization of large-scale AI inference systems.

This is a high-impact role at the intersection of machine learning research and distributed systems engineering, where you will drive how frontier models are deployed, scaled, and experienced in production.

You will play a key role in shaping our inference architecture, pushing the limits of latency, throughput, and cost efficiency, and translating cutting-edge research into robust, real-world systems.


Key Responsibilities

  • Lead the design and implementation of high-performance inference systems for LLMs, multimodal, and speech models
  • Drive improvements in:
  • Latency (P50 / P99)
  • Throughput (tokens/sec, QPS)
  • Cost efficiency (per request / per token)
  • Architect and optimize distributed inference systems across GPU clusters
  • Own and implement advanced techniques such as:
  • Quantization (INT8, FP8, AWQ, GPTQ)
  • Speculative decoding
  • KV-cache optimization and memory management
  • Model parallelism (tensor, pipeline, MoE routing)
  • Evaluate and integrate serving frameworks (vLLM, TensorRT-LLM, Triton, custom runtimes)
  • Partner with research teams to productionize new models and architectures
  • Build and maintain benchmarking and evaluation pipelines for inference performance
  • Diagnose and resolve bottlenecks across:
  • Model execution
  • GPU utilization
  • Networking and distributed systems
  • Mentor engineers and contribute to technical direction and best practices


Key Skills and Qualifications:


Core Requirements

  • 3–5+ years of experience in ML systems, backend engineering, or high-performance computing
  • Strong expertise in:
  • Deep learning frameworks (PyTorch, JAX, TensorFlow)
  • Python and/or C++
  • Deep understanding of:
  • Transformer architectures and modern LLM systems
  • GPU architecture and parallel computing
  • Distributed systems design

Experience

  • Proven track record building or optimizing large-scale inference systems in production
  • Hands-on experience with:
  • Model serving frameworks (vLLM, Triton, TensorRT-LLM, Ray Serve, etc.)
  • GPU optimization (CUDA, kernel tuning, memory management)
  • Model compression techniques (quantization, pruning, distillation)
  • Experience scaling systems handling high concurrency workloads

Nice to Have

  • Experience with compiler stacks (XLA, TVM, MLIR)
  • Familiarity with hardware accelerators (NVIDIA, AMD, TPUs)
  • Contributions to open-source ML systems
  • Background in multimodal or speech model serving
  • Published research in ML systems or efficiency


What Sets You Apart

  • You instinctively think in:
  • tokens/sec, GPU utilization, and tail latency
  • You can bridge:
  • Research ideas → production systems
  • You are comfortable operating across the stack:
  • Model → runtime → infrastructure
  • You take ownership of performance as a product feature


Leadership Expectations

  • Own critical parts of the inference stack and roadmap
  • Drive technical decisions and trade-offs across teams
  • Mentor junior engineers and elevate team standards
  • Influence system design across model, infra, and product layers


Why Join Us

  • Work on frontier AI systems deployed at scale
  • Solve deeply technical challenges in efficiency and systems design
  • Be part of building sovereign AI infrastructure
  • Shape how AI reaches millions of users in real-world applications
Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search