Back to search
Hydra Host Linkedin · Posted today

Support Engineer, GPU Infrastructure

Miami, Florida, United States

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Support Engineer, GPU Infrastructure (Tier 2/3)

Full-time | Support Operations | Reports to the Operations and Support Lead

Location: Remote

About Hydra Host

Hydra Host sells production-ready bare metal GPU compute for AI training, inference, enterprise workloads, and high-performance computing. Our control plane, Brokkr, connects customers to capacity across a network of partner-owned data centers rather than facilities we own ourselves.

When a customer's workload degrades, the cause sits somewhere across their code, our platform, the facility, or the hardware vendor. Most support roles decide what is broken. Here you also decide whose it is, on live production hardware, with a customer waiting and a partner relationship on the other side of the answer.

The role

You are the person a customer issue reaches when it is real. You take it from the first symptom to a resolution that holds, working from the Linux host down through the hardware and out to the facility floor.

The tier in the title is deliberate. Nearly everything that arrives at support here is already a Tier 2 or Tier 3 problem, on production hardware, with a customer's workload affected.

Most of the day is diagnosis, coordinating the people with hands on the hardware, and writing down what you found so the next person does not start over. Some of it crosses into infrastructure engineering work.

Support at Hydra is being built right now and you are one of the first hires into it. Expect less structure than you are used to, and more influence over what the structure becomes.

What You Will Do

Diagnose across the stack

  • Work Linux server issues end to end. Boot and network boot failures, kernel and driver problems, filesystems, storage pressure, services, memory and CPU behavior, general instability.
  • Diagnose hardware failures using out-of-band management, sensor data, POST and boot errors, SMART data, and vendor diagnostics. IPMI, Redfish, iDRAC, iLO, or equivalent.
  • Isolate server-side network problems. NICs and drivers, VLANs, addressing, routing, MTU, DNS, DHCP, bonding, link state, and packet captures when it comes to that.
  • Troubleshoot NVIDIA GPU servers. GPU availability, thermal throttling, driver and VBIOS mismatch, PCIe, XID errors, and the host-level conditions that look like GPU problems and are not.
  • Separate hardware from OS from network from application from configuration before escalating. Being wrong about which one it is costs more than being slow.

Own the incident, not just the ticket

  • Take an issue through to resolution or to a clean handoff, including the last step. A node that is repaired but still cordoned is not resolved.
  • Set severity by blast radius and communicate it. One node and one cluster are different events.
  • Notice when several tickets are one problem. Adjacent addresses, shared symptom, same time window.
  • Escalate to engineering with an evidence pack rather than a description, so the receiving engineer starts from your work instead of repeating it.
  • Take part in root cause analysis and post-incident review, and turn the findings into something that changes.

Work the partner and vendor boundary

  • Drive issues with data center partners. Remote hands, reboots, cabling and optics checks, component replacement, physical inspection.
  • Open and track hardware RMAs with OEMs through to a replacement in the rack.
  • Validate repaired or replaced equipment before it returns to production.
  • Keep the asset record accurate. When the register and the floor disagree, close the gap rather than working around it.
  • Support server turn-ups, migrations and decommissions where support is involved, and validate readiness before a machine carries a customer workload.

Communicate with customers

  • Write clear updates to technically sophisticated customers who want cause and timeline, not reassurance.
  • Ask for diagnostic information in a way that never implies the customer has misread their own situation.
  • Deliver an unwelcome answer plainly when that is the honest one.

Document What You Learn

  • Write the runbook after you solve something the first time, not the fifth.
  • Use accurate categories and real closure reasons. This data is how the team learns what is actually breaking.
  • Improve the runbooks and operational procedures you inherit as you use them.

Automation and continuous improvement

  • Build scripts and small tools that take repetitive diagnostic and support work off the queue.
  • Use Python, Bash, or similar to automate health checks, data collection, and routine operations.
  • Improve the troubleshooting tools and workflows you inherit rather than working around them.
  • Convert recurring manual procedures into documented ones, then into automated ones.
  • Contribute to infrastructure-as-code and configuration management where it touches support work.
  • Help improve monitoring and alerting. Alerts that fire in the wrong place, or do not fire at all, are a support problem before they are anyone else's.
  • Work with engineering to find the changes that make the platform easier to operate and cheaper to support at scale.

Coverage

Support runs across time zones and this role works a set schedule.

  • A defined shift, agreed before you start.
  • An escalation rotation for high severity issues outside your shift hours.
  • A written handoff at the end of every shift. What is open, what you tried, what to pick up first.

Your time zone matters to this hire. We will be direct in the first conversation about the hours we need covered.

What We Are Looking For

  • Three or more years supporting production servers, data center infrastructure, or bare metal and cloud environments.
  • Strong hands-on Linux troubleshooting. You are comfortable on a console with logs, dmesg, systemd, storage tooling, and network utilities.
  • Real experience with server hardware. CPU and memory, storage and filesystems, RAID, PCIe, NICs, power, BIOS and UEFI, firmware and drivers.
  • Out-of-band management experience. IPMI, Redfish, iDRAC, iLO, or similar.
  • Working TCP/IP knowledge and the ability to prove whether a problem is on the host or on the network.
  • You reason in fault domains. How much is broken, does it survive a rebuild, does it follow the workload to another machine, can it be fixed remotely.
  • Experience working in ticketing, monitoring, incident management, or infrastructure management systems.
  • Clear written English. Most of this job happens in writing, to customers, to partners, and to engineers.
  • Sound judgment alone in production. You will make calls at hours when nobody is available to check them.

Helpful, Not Required

  • NVIDIA GPU servers at scale, and any of CUDA, NCCL, NVLink, or DCGM.
  • HPC or AI training environments. InfiniBand or high performance Ethernet.
  • Enterprise platforms from Dell, HPE, Supermicro, or Lenovo.
  • NVMe, ZFS, Ceph, or distributed storage.
  • Prometheus, Grafana, or similar observability tooling.
  • NetBox or another infrastructure and asset register.
  • Ansible, Terraform, or configuration management. Git-based infrastructure workflows.
  • Optics, transceivers, and DAC or AOC cabling.
  • Working across geographically distributed third-party facilities.

What Success Looks Like

By 90 days

  • Working the majority of your shift's issues without escalation.
  • Every ticket you close carries an accurate category and a real closure reason.
  • At least three runbooks written from issues you personally resolved.
  • You know how to escalate to each facility contact on your shift without asking.

By six months

  • Your escalations to engineering arrive complete and are rarely handed back.
  • Issues on your shift are increasingly caught before the customer reports them.
  • Something that used to be a recurring ticket is gone because you removed the cause.
Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search