Senior AI Operator Optimization Engineer
Indexed description
Morgan McKinley is delighted to be partnering with a leading telecom technology organisation to hire for a Senior AI Operator Optimization Engineer role.
Please note this is a 12 months contract initially with annual review for extension.
This role is open for both onsite work in Dublin and remote work from other parts of EU.
The organisation has a Research Centre based in Dublin and with an ambition to expand their work in hardware, they are seeking a highly motivated AI Operator/Kernel Optimization Engineer to join their AI Infrastructure Platform team. In this role, you will focus on optimizing AI operators/kernels across diverse hardware architectures, driving high-performance execution for large-scale AI training and inference workloads. You will work closely with HQ computing software platform architects, compiler engineers, and AI researchers to maximize performance, scalability, and efficiency of AI computing systems.
They are a big believe of open-source contributions as one of the strongest indicators of engineering excellence and would be particularly interested in candidates who have evidence of their contributions to open-source AI communities.
Responsibilities
- Design, optimize, and accelerate AI operators/kernels on modern computing architectures, including CPUs, NPUs, AI accelerators, and heterogeneous computing platforms.
- Analyze hardware bottlenecks and optimize operator implementations to improve throughput, latency, memory efficiency, and power efficiency.
- Collaborate with computing software platform architecture teams to co-design next-generation AI computing systems software stacks.
- Develop architecture-aware optimization AI agents for deep learning operators, including compute-intensive and memory-intensive kernels.
- Optimize communication operators for distributed AI training and inference across large-scale computing clusters.
- Improve collective communication performance (e.g., AllReduce, AllGather, ReduceScatter, Broadcast) by optimizing communication algorithms, topology awareness, and overlap between communication and computation.
- Collaborate with compiler and runtime teams to enable automatic operator optimization and hardware-specific code generation.
- Benchmark, profile, and analyze AI workloads to identify optimization opportunities across the software-hardware stack.
- Stay up to date with emerging AI hardware architectures, compiler technologies, and distributed AI systems.
Qualifications
Required
- Master's degree or Ph.D. in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
- Strong experience in computer architecture, hardware architecture design, or hardware performance optimization.
- Hands-on experience optimizing AI operators or high-performance computing kernels on modern hardware platforms.
- Strong understanding of processor architecture, memory hierarchy, cache systems, SIMD/SIMT execution, and parallel computing.
- Experience with AI software frameworks such as vLLM, PyTorch, JAX, or ONNX Runtime.
- Proficiency in Triton, C/C++ and Python.
- Experience with GPU programming (CUDA, ROCm) or AI accelerator software stacks.
- Strong performance analysis and profiling skills.
Preferred
- Experience optimizing communication operators for distributed AI training.
- Experience with collective communication libraries such as NCCL, RCCL, MPI, Gloo, UCX, or oneCCL.
- Experience with AI compilers such as MLIR, TVM, XLA, Triton, LLVM, or Apache IREE.
- Experience with AI accelerator architectures (Ascend, TPU, Graphcore IPU, Cerebras, Tenstorrent, etc.).
- Familiarity with distributed systems, RDMA, InfiniBand, NVLink, PCIe, CXL, Ethernet fabrics, or high-performance networking.
- Publications or open-source contributions related to AI systems, HPC, compilers, or distributed AI.
- Strong understanding of AI system software, compilers, runtime systems, and hardware-software co-design.
- Experience in AI models performance modeling and bottleneck analysis.
- Excellent problem-solving and analytical skills.
- Strong communication and cross-functional collaboration abilities.
- Passion for pushing the performance limits of next-generation AI infrastructure.
Open Source Community Experience (preferred)
- Proven contributions to major AI open-source projects through merged pull requests (PRs), accepted code commits, feature implementations, performance optimizations, or bug fixes.
- Demonstrable evidence of code contributions (e.g., GitHub/GitLab contribution history, merged PRs, commit records, technical proposals, or release notes).
- Experience contributing to one or more leading AI open-source communities, Candidates with contributions to one or more of the following ecosystems are highly preferred: Triton, PyTorch, MLIR / LLVM, Hugging Face Transformers, NCCL / RCCL, JAX
- Experience serving as a Maintainer, Core Contributor, Reviewer, Technical Steering Committee (TSC) member, Module Owner, or Community Leader in a major open-source project is highly preferred.
- Active participation in open-source community governance, technical discussions, RFC reviews, or ecosystem development is considered a strong advantage.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search