Research Engineer - AI Workload & Systems
Indexed description
One of the goals of this lab are to enhance algorithm performance and training efficiency across industries, fostering long-term competitiveness.
About the job:
Frontier AI Technology Research
- Track the evolution of state-of-the-art AI model architectures, including Large Language Models (LLMs), Vision Language Models (VLMs), advanced attention mechanisms, and Mixture-of-Experts (MoE) architectures.
- Analyze the computational characteristics of emerging model architectures and Agentic AI training and inference workloads.
- Develop a systematic framework to map AI applications and workloads to hardware requirements.
- Identify key computational patterns, design efficient inference deployment strategies, and build analytical performance models.
- Evaluate the impact of algorithmic and hardware innovations on system performance, and provide quantitative, explainable architectural recommendations for next-generation AI accelerators.
Job requirements
About The Ideal Candidate:
- Strong understanding of modern AI model architectures and emerging trends, including sparse attention, linear attention, Mixture-of-Experts (MoE), and related techniques.
- Solid knowledge of AI accelerator architectures (e.g., GPUs, TPUs), with deep understanding of memory hierarchy, interconnect technologies, and hardware performance bottlenecks.
- Hands-on experience with AI kernel development using technologies such as Triton, TileLang, CUDA, or equivalent. Strong understanding of FlashAttention and other state-of-the-art kernel optimization techniques.
- Experience with modern LLM inference systems and optimizations, including vLLM, SGLang, or similar serving frameworks. Familiarity with the internals of deep learning frameworks such as PyTorch and JAX.
- Experience using hardware performance analysis and profiling tools, with the ability to develop Roofline-based performance analysis methodologies and tools.
- Ph.D. in Artificial Intelligence, Computer Architecture, Computer Systems, or a closely related field is an asset.
- Demonstrated research contributions through publications or influential open-source projects in AI infrastructure, systems, or computer architecture is an asset.
- Experience deploying and optimizing large-scale AI training or inference systems in production environments is an asset.
All applications for this position are reviewed directly by our hiring team, we do not use artificial intelligence tools to screen or select candidates.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search