Lead HPC Systems & Performance Engineer
Indexed description
Lead HPC Systems & Performance Engineer (Dallas, Texas):
The Company:
My client is a rapidly growing HPC and AI Supercompute Lab operating at the forefront of hyperscale compute.
With significant investment behind its growth, the business is building next-generation infrastructure supporting demanding AI, scientific research, simulation, and data-intensive workloads.
The Role:
They are looking for a Lead HPC Systems & Performance Engineer to take ownership of compute performance across large-scale HPC environments.
You'll evaluate emerging CPU, GPU and accelerator technologies, benchmark new hardware, identify performance bottlenecks, and turn real-world test data into architecture decisions.
Working across the compute stack, you'll collaborate with storage and networking specialists, hardware vendors, internal engineering teams and customers to build and optimise end-to-end HPC platforms.
Required Skills:
- Masters &/or PhD in Computer Science, Systems Engineering, or similar subject(s).
- Strong background in HPC systems, performance engineering or systems architecture.
- Deep understanding of CPU, GPU and accelerator architectures.
- Strong knowledge of memory hierarchy, NUMA and system performance.
- Hands-on experience with Linux tuning, optimisation and performance profiling.
- Experience designing or optimising HPC clusters and parallel computing environments.
- Understanding of InfiniBand/RoCE and their impact on compute performance.
- Experience benchmarking and evaluating new hardware.
- Knowledge of distributed/parallel storage and its relationship with compute performance.
- Experience supporting demanding AI/ML, scientific computing or simulation workloads.
Preferable:
- NVIDIA GPU ecosystem and tools such as DCGM, Nsight or MLPerf.
- Experience with large-scale GPU/CPU/DPU clusters & emerging/pre-production hardware.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search