Technical Lead - HPC
Indexed description
Role Summary
This is a hands-on Technical Lead who can support the build, operation and scale Era4’s sovereign AI/HPC infrastructure. We need someone who can operate at the intersection of HPC platform engineering, GPU infrastructure, Linux systems, Kubernetes/Slurm, automation, observability and production incident response.
You will act as a technical authority, helping shape how GPU clusters are deployed, monitored, automated, supported and improved. You will work closely with SRE, platform engineering, infrastructure, vendors and customer-facing teams to ensure our platform is reliable, observable, scalable and ready for production workloads.
Key Responsibilities
HPC & GPU Platform Leadership:
- Lead infrastructure, supporting GPU, compute, storage and networking platforms.
- Technical authority across platform engineering, SRE, infrastructure and customer-facing teams.
- Mentor engineers and help define technical standards, best practices and operational excellence.
- Lead technical investigations during major incidents and production outages.
- Improve observability, monitoring and alerting across the platform.
- Drive root-cause analysis and implement long-term reliability improvements.
- Improve platform scalability, deployment processes and operational efficiency.
- Oversee runbook creation, automation and development of inhouse Agent capabilities to support Operations
- Contribute to the design, build, enhancement of future capabilities
- Support customer onboarding, complex technical escalations and platform adoption.
- Work with vendors, partners and internal teams to resolve infrastructure issues.
- Translate complex technical challenges into clear, actionable communication.
- Contribute to Technical and Operational Roadmaps
- Contribute to the deployment and optimisation of GPU clusters, Kubernetes environments and next-generation AI infrastructure.
- Linux engineering background
- Experience in production of leading teams supporting HPC, GPU infrastructure, AI infrastructure, research computing, cloud HPC, Neocloud, platform engineering or SRE.
- Experience operating or supporting production infrastructure across compute, networking, storage and observability.
- Hands-on experience with at least one workload or orchestration layer such as Slurm, Kubernetes, Run:ai, LSF, PBS or equivalent.
- Experience with Open Source monitoring and troubleshooting using tools such as Prometheus, Grafana, OpenTelemetry, Loki or equivalent.
- Proven involvement in major incidents, on-call, escalation, root-cause analysis or production troubleshooting.
- Ability to mentor engineers, influence technical direction and lead by technical credibility.
- Comfortable working with internal teams, customers, suppliers and vendor engineering teams.
Diversity & Inclusion
Era4 is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search