Back to search
Acceler8 Talent Linkedin · Posted 1mo ago

Infrastructure Engineer

San Francisco, California, United States

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Infrastructure Platform Engineer (AI Infrastructure)

San Francisco, CA (Onsite / Flex to 3 days in office)

$200,000-$350,000 Base + Equity



I'm working with a rapidly growing AI infrastructure company building the orchestration layer for the future of AI compute.


As AI infrastructure becomes increasingly heterogeneous, operating large fleets of GPUs is no longer enough. This team is building the platform that brings together CPUs, GPUs, and emerging accelerators into a single production inference cloud, allowing AI workloads to run wherever they're most efficient.


They've recently emerged from stealth with significant funding, eight-figure revenue, Fortune 500 deployments, and a growing roster of AI-native customers.


This isn't a traditional cloud infrastructure role.

You'll own the cluster infrastructure powering a heterogeneous AI cloud, working across bare metal, Linux, Kubernetes, high-speed networking, observability, and production operations. You'll partner closely with compiler, runtime, distributed systems, and hardware teams to ensure new hardware platforms can reliably serve production AI workloads from day one.


You'll work on problems such as:

• Operating large-scale CPU, GPU, and accelerator clusters for production AI inference

• Bare-metal provisioning, hardware lifecycle management, and cluster automation

• Kubernetes, Slurm, Nomad, and cluster scheduling systems

• High-performance networking across InfiniBand, RDMA, and modern datacenter fabrics

• Linux systems debugging across networking, storage, firmware, drivers, and kernel layers

• Fleet observability, capacity planning, incident response, and production reliability

• Bringing new accelerator platforms online and validating production readiness


We're looking for engineers who have:

• Owned production cluster infrastructure, from provisioning through day-to-day operations and incident response

• Deep Linux systems experience and confidence debugging networking, storage, drivers, firmware, and kernel-level issues

• Experience operating Kubernetes, Slurm, Nomad, or other production cluster schedulers

• Built automation for provisioning, upgrades, lifecycle management, or infrastructure operations

• Experience supporting GPU, accelerator, HPC, or AI infrastructure in production environments

• A strong operational mindset with a focus on reliability, observability, and automation.



Particularly relevant experience includes:

• Bare-metal provisioning (PXE/iPXE, image pipelines, rack bring-up)

• GPU cluster operations

• Kubernetes platform engineering

• InfiniBand, RDMA, or RoCE networking

• Fleet observability (Prometheus, Grafana, OpenTelemetry)

• Terraform, Ansible, Helm, Go, or Python

• CUDA / ROCm stacks

• Multi-tenant scheduling and resource management

• Hardware validation and firmware management

• AI inference or HPC infrastructure





If you're excited about owning the infrastructure that powers the next generation of AI compute, This might be worth a chat!

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search