Staff Software Engineer, SIMD Kernels
Indexed description
Location:
Santa Clara, CA headquarters, or any of our regional offices. Will consider remote.
The Role: Staff Software Engineer, SIMD Kernels
What You Will Do
As a member of the SIMD Kernels team, you will help productize the software stack for our AI compute engine. You will develop, enhance, and maintain software kernels for ML operators—such as softmax, layer norm, and activation functions—for next‑generation AI hardware. You will also build solutions that make our SDK intuitive for developers to use and analyze performance.
You have experience building software kernels for modern hardware architectures and understand how to map algorithms and AI‑framework computational graphs onto those architectures. You’re comfortable working across the full toolchain and navigating hardware–software co‑design trade‑offs. You can deliver and scale high‑quality software quickly in a fast‑moving development environment.
Minimum
What you will bring:
- MS or PhD in computer engineering, math, physics, or a related degree with 5+ years of industry experience.
- Strong grasp of computer architecture, data structures, system software, and machine learning fundamentals.
- Proficient in C/C++ and Python development in Linux environment and using standard development tools.
- Experience in implementing algorithms in C/C++ and Python.
- Experience in implementing algorithms for specialized hardware such as FPGAs, DSPs, GPUs, AI accelerators using libraries such as CUDA, etc.
- Experience in implementing operators commonly used in ML workloads—GEMMs, Convolutions, softmax, layer normalization, pooling, etc.
- Self-motivated team player with a strong sense of ownership and leadership.
- Prior startup, small team, or incubation experience.
- Experience with ML frameworks such as TensorFlow and/or PyTorch.
- Experience working with ML compilers and algorithms, such as MLIR, LLVM, TVM, Glow, etc.
- Experience with a deep learning framework (such as PyTorch or TensorFlow) and ML models for CV, NLP, or recommendation.
- Experiences in development for embedded SIMD vector processors such as Tensilica.
- Work experience at a cloud provider or AI compute/subsystem company.
d-Matrix does not accept resumes or candidate submissions from external agencies. We appreciate the interest and effort of recruitment firms, but we kindly request that individual interested in opportunities with d-Matrix apply directly through our official channels. This approach allows us to streamline our hiring processes and maintain a consistent and fair evaluation of al applicants. Thank you for your understanding and cooperation.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search