Back to search
Insight Global Linkedin · Posted 3d ago

Compute Engineer

New York City, New York, United States

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

COMPANY OVERVIEW

A leading AI infrastructure company focused on building and operating large-scale compute environments that power next-generation AI workloads. The organization designs, builds, and operates data center infrastructure at massive scale with teams spanning both hardware and software. Speed, scale, and operational excellence are core differentiators, with a focus on delivering frontier compute infrastructure for AI.


HOW THEY OPERATE

Extreme ownership. Full autonomy. Own things end to end often taking on scope outside your core role without being asked to get things done.

Velocity. We drive everything forward as fast as possible.

First principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.

Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.


THE INFRASTRUCTURE TEAM

Examples of key problems the team is working on

Bring gigawatts of accelerators from first power-on to production. Facility availability to ready-for-service across thousands of racks per site, with a new data hall landing every few weeks.

Make rack qualification faster than the fleet grows. Firmware baselines, burn-in, and cluster validation proven on every rack before a customer workload touches it, at a pace that never becomes the critical path.

Scale by tooling, not headcount. Deployed megawatts grow severalfold next year while the team stays near-flat, because anything done twice by hand becomes software.


ROLE SCOPE

  • Own compute turn-up from facility availability to ready-for-service: the stretch after the network hands off and before customers run workloads.
  • Qualify racks at scale: establish firmware baselines, configure BMC and BIOS, run burn-in, and validate at node and cluster level across hundreds of racks per site on GPU and custom accelerator platforms.
  • Drive qualification through the base-management Kubernetes platform and provisioning stack (discovery, imaging, firmware updates, shared services), burning down qual queues with tooling rather than manual runs.
  • Triage hardware failures found in qualification: isolate to component, drive RMA and vendor escalation, and feed failure patterns back into the qual gates.
  • Run turn-up remotely by default, with on-site pulses of roughly a week per data hall as new halls reach facility availability, plus occasional overlapping-site weeks.
  • Partner with network deployment, ICT, data center operations, and hardware teams during turn-up windows, and support incident response on freshly-live capacity.
  • Ability to travel 20-30% of the time to Data Centers and Labs, as needed.


WHAT WE'RE LOOKING FOR

The below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly, we'd still welcome a conversation.

Must Haves

  • You've brought up server or GPU fleets at scale, hundreds of nodes or more, and taken them all the way to production.
  • You work deep in Linux and out-of-band management: BMC, IPMI, and Redfish are daily tools for you, not occasional lookups.
  • You've automated hardware workflows in Python or Go rather than clicking through them, and the second time you do anything by hand you turn it into software.
  • You've worked physically in data halls, racking, cabling, and swapping components, and you're just as effective acting as remote hands or directing them.
  • You triage failures methodically across hardware, firmware, and software, isolating the fault to a component before reaching for a fix.
  • You travel for turn-up windows when a new data hall comes online.

Preferred

  • Kubernetes-based bare-metal provisioning.
  • Accelerator platform bringup (NVIDIA, AMD, or custom).
  • Burn-in and stress harness design.
  • DCIM and inventory tooling.


COMPENSATION

$164K - $206K Base + Equity (depending on experience)


IDEAL CANDIDATE

Linux-heavy infrastructure engineer with experience bringing server or GPU infrastructure from deployment through production readiness. Strong background in hardware qualification, automation, firmware management, troubleshooting, and operating within hyperscale or large-scale data center environments.

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search