Staff Engineer, Inference Optimizations
Indexed description
DigitalOcean Staff Engineer, Inference Optimizations Reposted 12 Hours AgoSaved In-Office Denver, CO, USA 191K-239K Annually Senior level 191K-239K Annually Senior levelArtificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)Lead architecture and low-level optimizations for high-performance AI inference. Benchmark and tune GPU kernels, memory and precision management, parallelize across multi-node GPU clusters, implement quantization (FP8/INT8/FP4), advise on hardware/software stacks (CUDA/ROCm/Triton/TensorRT), mentor engineers, and collaborate to turn hardware limits into shippable inference products.Top Skills: Amd AiterAsk (Assembly)Bf16Ck (Composable Kernel)CudaFlashattentionFp4Fp8Int8Openai TritonRocmTensorrtTriton Compiler
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search