Infrastructure Engineer
Indexed description
Infrastructure Platform Engineer (AI Infrastructure)
San Francisco, CA (Onsite / Flex to 3 days in office)
$200,000-$350,000 Base + Equity
I'm working with a rapidly growing AI infrastructure company building the orchestration layer for the future of AI compute.
As AI infrastructure becomes increasingly heterogeneous, operating large fleets of GPUs is no longer enough. This team is building the platform that brings together CPUs, GPUs, and emerging accelerators into a single production inference cloud, allowing AI workloads to run wherever they're most efficient.
They've recently emerged from stealth with significant funding, eight-figure revenue, Fortune 500 deployments, and a growing roster of AI-native customers.
This isn't a traditional cloud infrastructure role.
You'll own the cluster infrastructure powering a heterogeneous AI cloud, working across bare metal, Linux, Kubernetes, high-speed networking, observability, and production operations. You'll partner closely with compiler, runtime, distributed systems, and hardware teams to ensure new hardware platforms can reliably serve production AI workloads from day one.
You'll work on problems such as:
• Operating large-scale CPU, GPU, and accelerator clusters for production AI inference
• Bare-metal provisioning, hardware lifecycle management, and cluster automation
• Kubernetes, Slurm, Nomad, and cluster scheduling systems
• High-performance networking across InfiniBand, RDMA, and modern datacenter fabrics
• Linux systems debugging across networking, storage, firmware, drivers, and kernel layers
• Fleet observability, capacity planning, incident response, and production reliability
• Bringing new accelerator platforms online and validating production readiness
We're looking for engineers who have:
• Owned production cluster infrastructure, from provisioning through day-to-day operations and incident response
• Deep Linux systems experience and confidence debugging networking, storage, drivers, firmware, and kernel-level issues
• Experience operating Kubernetes, Slurm, Nomad, or other production cluster schedulers
• Built automation for provisioning, upgrades, lifecycle management, or infrastructure operations
• Experience supporting GPU, accelerator, HPC, or AI infrastructure in production environments
• A strong operational mindset with a focus on reliability, observability, and automation.
Particularly relevant experience includes:
• Bare-metal provisioning (PXE/iPXE, image pipelines, rack bring-up)
• GPU cluster operations
• Kubernetes platform engineering
• InfiniBand, RDMA, or RoCE networking
• Fleet observability (Prometheus, Grafana, OpenTelemetry)
• Terraform, Ansible, Helm, Go, or Python
• CUDA / ROCm stacks
• Multi-tenant scheduling and resource management
• Hardware validation and firmware management
• AI inference or HPC infrastructure
If you're excited about owning the infrastructure that powers the next generation of AI compute, This might be worth a chat!
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search