Data Center Engineer — Kubernetes
Indexed description
Infrastructure & Deployment
Data Center Engineer — Kubernetes
Design and operate the infrastructure that connects physical data-center machines, AWS EKS, Kubernetes clusters, and PERCO's worker system into one reliable deployment platform.
Applications open
Location Remote / San Francisco, CA
Work model Remote-first
Employment Full-time
Compensation Salary range shared in process
Application window Open until filled
Apply for this role ↓
01 / The role
Build the system, then prove it works.
CaelumenAI is building the governed production layer for agent-native software. PERCO needs an infrastructure foundation that can deploy and operate workloads across public cloud and customer-controlled data centers without losing security, observability, or human control. As a Data Center Engineer — Kubernetes, you will help design data-center solutions, bring physical machines into service, operate Kubernetes on AWS EKS and self-managed environments, and connect those environments to PERCO's worker system. You will own the path from capacity and network planning through provisioning, cluster lifecycle, production diagnostics, and verified recovery. This is a remote-first role based in the United States with an option to work from San Francisco. Planned travel to data-center or customer sites is required when physical deployment, commissioning, or incident response calls for it. Compensation is discussed during the interview process. Applicants must have valid authorization to work in the United States when employment begins. We welcome candidates working under valid OPT and support eligible H-1B sponsorship.
02 / Your work
What you’ll do
- Design practical data-center deployment solutions across compute, storage, networking, power, access, and operational constraints
- Provision, commission, inventory, and troubleshoot physical Linux servers and their supporting network paths
- Build and operate Kubernetes clusters across AWS EKS and self-managed infrastructure
- Integrate clusters and machines with PERCO's worker system for controlled scheduling, execution, status reporting, and lifecycle management
- Automate repeatable provisioning, upgrades, scaling, recovery, and compliance checks with infrastructure as code
- Establish observability for cluster health, workloads, capacity, networking, and deployment evidence
- Diagnose production failures across hardware, Linux, networking, Kubernetes, cloud services, and application workloads
- Document runbooks and leave every incident with a root-cause fix or a clearly owned follow-up
What we’re looking for
Evidence over pedigree
Show us something you scoped, built, tested, and debugged—and tell us what the evidence changed.
- Production experience operating Kubernetes and Linux infrastructure
- Hands-on experience with AWS and EKS, including networking, identity, storage, and cluster lifecycle
- Ability to provision and troubleshoot physical or virtual machines, networking, DNS, certificates, and container runtimes
- Experience automating infrastructure with tools such as Terraform, Ansible, Helm, or equivalent systems
- Strong incident-debugging habits across logs, metrics, events, and system state
- Ability to communicate architecture decisions, operating risks, and recovery evidence clearly
- Willingness and ability to travel to U.S. data-center or customer sites when planned work requires it
Useful, not required
- Experience with bare-metal Kubernetes, Cluster API, Talos Linux, PXE or image-based provisioning, or immutable host systems
- Experience with GPU infrastructure, workload scheduling, or distributed worker systems
- Experience with hybrid-cloud networking, VPNs, private connectivity, or multi-cluster operations
- Experience designing secure customer-managed or regulated deployment environments
What you get
- Ownership of a hybrid infrastructure platform spanning cloud and data centers
- Direct influence on PERCO's machine and worker-system architecture
- Flexible remote collaboration with an option to work from San Francisco
- Planned, purposeful travel rather than continuous on-site staffing
- Support for candidates working under valid OPT and eligible H-1B sponsorship
Apply for this role
Tell us what you have built and how you know it worked. A concise application with real evidence is enough.
Company website
Full name Email
Phone Optional Location Optional
LinkedIn Optional Portfolio or GitHub Optional What should we know? Optional Upload resume PDF, DOC, DOCX, or TXT Upload cover letter OptionalPDF, DOC, DOCX, or TXT
Anti-bot verification protects this form.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search