AI Kernel Writer
Indexed description
We’re a fast-moving AI startup building next-generation infrastructure for the world’s most demanding AI workloads. Our mission is to accelerate the future of intelligence by delivering cutting-edge server solutions optimized for large-scale inference and training. Backed by top-tier investors and led by industry veterans, we’re scaling rapidly and looking for a dynamic partnerships leader to help us win the trust of enterprise and hyperscaler customers.
We are a US-Israeli company looking for builders-engineers and operators who have shipped production hardware and know the crucial difference between compelling research and systems that operate at massive scale.
If you want to build infrastructure that fundamentally changes what the world can accomplish with AI, come join us!
About The Position
- Design and implement high-performance compute kernels for AI primitives such as GEMM, attention, normalization, and convolution.
- Optimize for throughput, latency, and memory hierarchy across heterogeneous compute units (SIMD, matrix engines, DMA).
- Collaborate with compiler and runtime teams to integrate kernels into Triton and PyTorch pipelines.
- Profile and tune kernels using tools like Perfetto, VTune, Tracy, Nsight, or custom simulators.
- Prototype and evaluate precision formats (FP16/BF16/FP8/e5m2, MXFP/FP4, etc.).
- Contribute to micro-architecture feedback loops, helping co-design ISA and memory features with the hardware team.
- Degree at any level (Bachelor's, Master's, or Ph.D.) in Computer Science, Computer Engineering, or a related field from a recognized university.
- Strong background in parallel programming (CUDA, Triton, RISC-V Vector (RVV), POSIX Threads, or OpenMP).
- Deep understanding of memory layout, vectorization, thread/block scheduling, and cache behavior.
- Experience programming wide-SIMD vector units and matrix/systolic engines.
- Familiarity with explicit, DMA-based data movement and software-managed scratchpad memories.
- Familiarity with the parallel-computing ecosystem and its high-performance libraries and primitives, such as BLAS/BLIS, dense linear algebra, and parallel scan/reduce/sort.
- Skilled in performance analysis and parallel debugging using tools such as GNU Debugger and Nsight.
- Hands-on experience profiling and optimizing compute or AI workloads (e.g., GEMM, softmax, attention).
- Solid grasp of numerical stability, precision formats, and mixed precision arithmetic.
- Proficiency in C++17 or higher, with strong knowledge of standard algorithms, data structures, and generic programming paradigms.
- Collaborative work style with the ability to operate effectively in multicultural, cross-disciplinary environments.
Majestic is proudly an equal opportunity employer deeply committed to diversity of race, religion, gender, sexual orientation, background, and thought. If you are ready to dare greatly, learn from mistakes, bring your authentic voice to vigorous debates, and help us create a step function in human technical capability, you belong here.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search