Software Engineer - Cloud Infrastructure
Indexed description
Inference is an unforgiving workload for Kubernetes. Traffic is bursty and latency-sensitive, GPU capacity is scarce and inelastic, tenants must stay isolated, and multi-node serving depends on the network holding up under sustained load. This is a hands-on architecture role for an engineer who has already run large clusters in production and wants to push them further.
Key Responsibilities
Cluster Architecture
- Own the architecture of our multi-cluster, multi-tenant Kubernetes fleet across both managed and self-managed clusters: cluster topology, control plane and etcd lifecycle, and zero-downtime upgrades.
- Extend Kubernetes with custom controllers, operators, and CRDs so platform behavior is encoded in software rather than runbooks.
- Design GPU scheduling and capacity strategy, including topology-aware placement, node pools, priority and preemption, and quota across tenants.
- Build autoscaling that matches inference traffic: queue-driven pod scaling, node autoscaling, scale-to-zero, and cold-start reduction.
- Own the Kubernetes network data plane: CNI, IPAM, DNS, ingress, and L4/L7 load balancing.
- Design cross-AZ, cross-region, and cross-cluster connectivity, and operate the service mesh for routing, mTLS, and traffic policy.
- Debug production network issues (packet loss, conntrack exhaustion, MTU mismatches, DNS latency, load balancer behavior) and drive permanent fixes.
- Define SLOs for platform-critical systems and lead post-incident hardening.
- Deliver infrastructure as code with Terraform, Helm, and GitOps.
- Partner with the inference engine, platform, SRE, and security teams to turn serving requirements into platform capabilities.
- 5+ years designing, building, and operating large-scale Kubernetes infrastructure in production.
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent.
- Proven experience operating large-scale, high-traffic network services in production.
- Deep understanding of Kubernetes internals: API server, scheduler, controller loops, kubelet, and etcd.
- Strong command of Kubernetes and cloud networking: CNI, kube-proxy/eBPF datapaths, DNS, load balancing, service mesh, and VPC routing.
- Proficiency with AWS, Terraform, Helm, and Ansible.
- Programming skills in Go or Python, with the ability to build infrastructure tooling and automation.
- Strong debugging skills across distributed systems, containers, and the Linux networking stack.
- Clear written and verbal communication, including the ability to document architectural decisions for other engineers.
- Large-scale Kubernetes operations in a high-traffic domain such as gaming, e-commerce, or public cloud.
- Cilium and eBPF, including kube-proxy replacement or upstream contributions.
- Cluster provisioning and lifecycle management with Kubespray or similar Ansible-based tooling.
- GPU orchestration: NVIDIA GPU Operator, device plugins, or Dynamic Resource Allocation (DRA).
- High-performance networking for distributed workloads: RDMA/RoCE, InfiniBand, EFA, SR-IOV, or NCCL tuning.
- Multi-cloud, hybrid-cloud, or bare-metal Kubernetes operations.
- Contributions to Kubernetes, Cilium, Istio, or other CNCF projects.
- Flexible working hours
- Daily lunch and dinner provided; unlimited snacks and beverages
- Supportive and highly collaborative work environment
- Health check-up support and top-tier equipment/hardware support
- A front-row seat to the generative AI infrastructure revolution
- Competitive compensation, startup equity, health insurance, and other benefits.
We are a small, fast-moving team doing work that matters at one of the most exciting moments in the history of technology. With our world-class inference stack, we are building the platform teams can actually rely on.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search