Back to search
DigitalOcean Builtin · Indexed 2026-07-24

Staff Engineer, Inference Optimizations

Boston, MA, USA 191K-239K Annually

Senior level Builtin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

DigitalOcean Staff Engineer, Inference Optimizations YesterdaySaved In-Office Boston, MA, USA 191K-239K Annually Senior level 191K-239K Annually Senior levelArtificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)Lead architecture and low-level optimizations for LLM inference to maximize throughput and minimize latency. Profile and optimize GPU kernels, implement quantization and parallelization strategies (FP8/INT8/FP4), tune libraries (AITER, Triton, CUDA, ROCm), and mentor engineers while translating hardware limits into shippable platform features.Top Skills: Ai/Ml InferenceAmd AiterAmd Mi355XBf16CudaDeepseekFlashattentionFp4Fp8Glm-5Gpu Kernel DevelopmentInt8MoeMulti-Node Gpu ClusteringOpenai TritonQwen3-235BRmsnormRocmTensorrtTflopsTriton Compiler

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search