Staff Engineer, Inference Optimizations
Indexed description
DigitalOcean Staff Engineer, Inference Optimizations Reposted 20 Hours AgoSaved In-Office San Francisco, CA, USA 191K-239K Annually Senior level 191K-239K Annually Senior levelArtificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)Lead benchmarking and performance optimization for AI inference at GPU/kernel levels. Solve memory bandwidth and compute bottlenecks, implement quantization and parallelization (multi-node), advise on hardware/software stacks, mentor engineers, and collaborate to ship high-performance inference features.Top Skills: AiterAmd GpusAskBf16CkCudaCustom Cuda KernelsFlashattentionFp4Fp8Int8MoeNvidia GpusOpenai TritonRmsnormRocmTensorrtTritonTriton Compiler
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search