Staff Engineer, Inference Optimizations
Indexed description
DigitalOcean Staff Engineer, Inference Optimizations YesterdaySaved In-Office Boston, MA, USA 191K-239K Annually Senior level 191K-239K Annually Senior levelArtificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)Lead architecture and low-level optimizations for LLM inference to maximize throughput and minimize latency. Profile and optimize GPU kernels, implement quantization and parallelization strategies (FP8/INT8/FP4), tune libraries (AITER, Triton, CUDA, ROCm), and mentor engineers while translating hardware limits into shippable platform features.Top Skills: Ai/Ml InferenceAmd AiterAmd Mi355XBf16CudaDeepseekFlashattentionFp4Fp8Glm-5Gpu Kernel DevelopmentInt8MoeMulti-Node Gpu ClusteringOpenai TritonQwen3-235BRmsnormRocmTensorrtTflopsTriton Compiler
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search