Back to search
CaelumenAI Linkedin · Posted 1mo ago

Data Center Engineer — Kubernetes

San Francisco, CA, United States

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

← All roles

Infrastructure & Deployment

Data Center Engineer — Kubernetes

Design and operate the infrastructure that connects physical data-center machines, AWS EKS, Kubernetes clusters, and PERCO's worker system into one reliable deployment platform.

Applications open

Location Remote / San Francisco, CA

Work model Remote-first

Employment Full-time

Compensation Salary range shared in process

Application window Open until filled

Apply for this role ↓

01 / The role

Build the system, then prove it works.

CaelumenAI is building the governed production layer for agent-native software. PERCO needs an infrastructure foundation that can deploy and operate workloads across public cloud and customer-controlled data centers without losing security, observability, or human control. As a Data Center Engineer — Kubernetes, you will help design data-center solutions, bring physical machines into service, operate Kubernetes on AWS EKS and self-managed environments, and connect those environments to PERCO's worker system. You will own the path from capacity and network planning through provisioning, cluster lifecycle, production diagnostics, and verified recovery. This is a remote-first role based in the United States with an option to work from San Francisco. Planned travel to data-center or customer sites is required when physical deployment, commissioning, or incident response calls for it. Compensation is discussed during the interview process. Applicants must have valid authorization to work in the United States when employment begins. We welcome candidates working under valid OPT and support eligible H-1B sponsorship.

02 / Your work

What you’ll do

  • Design practical data-center deployment solutions across compute, storage, networking, power, access, and operational constraints
  • Provision, commission, inventory, and troubleshoot physical Linux servers and their supporting network paths
  • Build and operate Kubernetes clusters across AWS EKS and self-managed infrastructure
  • Integrate clusters and machines with PERCO's worker system for controlled scheduling, execution, status reporting, and lifecycle management
  • Automate repeatable provisioning, upgrades, scaling, recovery, and compliance checks with infrastructure as code
  • Establish observability for cluster health, workloads, capacity, networking, and deployment evidence
  • Diagnose production failures across hardware, Linux, networking, Kubernetes, cloud services, and application workloads
  • Document runbooks and leave every incident with a root-cause fix or a clearly owned follow-up

03 / The signal

What we’re looking for

Evidence over pedigree

Show us something you scoped, built, tested, and debugged—and tell us what the evidence changed.

  • Production experience operating Kubernetes and Linux infrastructure
  • Hands-on experience with AWS and EKS, including networking, identity, storage, and cluster lifecycle
  • Ability to provision and troubleshoot physical or virtual machines, networking, DNS, certificates, and container runtimes
  • Experience automating infrastructure with tools such as Terraform, Ansible, Helm, or equivalent systems
  • Strong incident-debugging habits across logs, metrics, events, and system state
  • Ability to communicate architecture decisions, operating risks, and recovery evidence clearly
  • Willingness and ability to travel to U.S. data-center or customer sites when planned work requires it

04 / Bonus signal

Useful, not required

  • Experience with bare-metal Kubernetes, Cluster API, Talos Linux, PXE or image-based provisioning, or immutable host systems
  • Experience with GPU infrastructure, workload scheduling, or distributed worker systems
  • Experience with hybrid-cloud networking, VPNs, private connectivity, or multi-cluster operations
  • Experience designing secure customer-managed or regulated deployment environments

05 / What we offer

What you get

  • Ownership of a hybrid infrastructure platform spanning cloud and data centers
  • Direct influence on PERCO's machine and worker-system architecture
  • Flexible remote collaboration with an option to work from San Francisco
  • Planned, purposeful travel rather than continuous on-site staffing
  • Support for candidates working under valid OPT and eligible H-1B sponsorship

Application

Apply for this role

Tell us what you have built and how you know it worked. A concise application with real evidence is enough.

Company website

Full name Email

Phone Optional Location Optional

LinkedIn Optional Portfolio or GitHub Optional What should we know? Optional Upload resume PDF, DOC, DOCX, or TXT Upload cover letter OptionalPDF, DOC, DOCX, or TXT

Anti-bot verification protects this form.

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search