Back to search
Stanford Black Limited Linkedin · Posted 20d ago

Lead HPC Systems & Performance Engineer

Dallas, Texas, United States

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Lead HPC Systems & Performance Engineer (Dallas, Texas):


The Company:

My client is a rapidly growing HPC and AI Supercompute Lab operating at the forefront of hyperscale compute.


With significant investment behind its growth, the business is building next-generation infrastructure supporting demanding AI, scientific research, simulation, and data-intensive workloads.


The Role:

They are looking for a Lead HPC Systems & Performance Engineer to take ownership of compute performance across large-scale HPC environments.


You'll evaluate emerging CPU, GPU and accelerator technologies, benchmark new hardware, identify performance bottlenecks, and turn real-world test data into architecture decisions.


Working across the compute stack, you'll collaborate with storage and networking specialists, hardware vendors, internal engineering teams and customers to build and optimise end-to-end HPC platforms.


Required Skills:

  • Masters &/or PhD in Computer Science, Systems Engineering, or similar subject(s).
  • Strong background in HPC systems, performance engineering or systems architecture.
  • Deep understanding of CPU, GPU and accelerator architectures.
  • Strong knowledge of memory hierarchy, NUMA and system performance.
  • Hands-on experience with Linux tuning, optimisation and performance profiling.
  • Experience designing or optimising HPC clusters and parallel computing environments.
  • Understanding of InfiniBand/RoCE and their impact on compute performance.
  • Experience benchmarking and evaluating new hardware.
  • Knowledge of distributed/parallel storage and its relationship with compute performance.
  • Experience supporting demanding AI/ML, scientific computing or simulation workloads.


Preferable:

  • NVIDIA GPU ecosystem and tools such as DCGM, Nsight or MLPerf.
  • Experience with large-scale GPU/CPU/DPU clusters & emerging/pre-production hardware.


Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search