Staff Engineer, Inference Optimizations
Indexed description
DigitalOcean Staff Engineer, Inference Optimizations Reposted 7 Hours AgoSaved In-Office Austin, TX, USA 191K-239K Annually Senior level 191K-239K Annually Senior levelArtificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)Lead benchmarking and low-level performance optimization for GPU-based inference. Solve memory, precision, and parallelization bottlenecks, tune kernels (CUDA/Triton/ROCm/AMD AITER), implement quantization (FP8/INT8/FP4), advise on hardware/software integration, mentor engineers, and translate hardware limits into shippable features for high-throughput, low-latency inference fleets.Top Skills: Amd Aiter (Ck/Ask)Amd GpusAmd Mi355XBf16CudaCuda KernelsFlashattentionFp4Fp8Int8Moe (Mixture Of Experts)Nvidia GpusOpenai TritonRms NormRocmTensorrtTriton Compiler
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search