HPC Network Engineer
Indexed description
Job Description
We are looking for an HPC Network Engineer to join our global team responsible for managing and supporting high-performance networking environments for large-scale AI infrastructure. You will help ensure the performance, reliability, and security of critical network fabrics, working closely with distributed compute and storage teams to support modern AI workloads.
This is an opportunity to work hands-on with some of the most advanced InfiniBand and Ethernet networking infrastructure in production today, while developing deep expertise in AI infrastructure and high-performance computing (HPC) technologies.
Responsibilities:
- Support the design, deployment, configuration, and maintenance of InfiniBand and Ethernet network infrastructures.
- Troubleshoot complex network issues, including connectivity, latency, routing, and performance degradation across hybrid environments.
- Manage and optimize high-performance network components, such as switches, Host Channel Adapters (HCAs), subnet managers, and fabric configurations.
- Implement, manage, and troubleshoot network security and firewall technologies (e.g., Fortinet solutions like FortiGate, VPNs).
- Monitor network health, perform performance tuning, and assist in capacity planning for HPC and AI networking systems.
- Collaborate with compute, storage, and platform teams to seamlessly support HPC and AI workloads.
- Participate in incident response, on-call support activities, and drive long-term operational improvements.
- Document network architectures, configurations, and operational procedures, and adopt automation tools (e.g., Ansible, Terraform) to streamline workflows.
- Proven experience in networking, system engineering, or data center infrastructure roles.
- Solid understanding of networking fundamentals, including TCP/IP, routing protocols (BGP, OSPF), switching, VLANs, QoS, and network design.
- Hands-on experience or strong familiarity with Linux operating systems and command-line tools.
- Effective verbal and written communication skills in English, with the ability to collaborate in a global team environment.
- Strong analytical and problem-solving skills to diagnose and resolve complex network challenges.
- Exposure to or hands-on experience with InfiniBand fabrics and high-performance networking concepts (e.g., NVIDIA/Mellanox).
- Familiarity with or certifications in enterprise firewall technologies (e.g., Fortinet FCSS/FCNSP, CCNP/CCIE).
- Experience with large-scale HPC clusters, AI/ML infrastructure, or distributed systems environments.
- Knowledge of RDMA, MPI, and low-latency networking concepts.
- Proficiency in scripting languages (e.g., Python, Bash) and Infrastructure-as-Code (IaC) tools for automation.
- Work with an established Silicon Valley leader in the cloud infrastructure industry;
- Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
- Be a part of cutting-edge, open-source innovation;
- Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
- Professional development and training;
- Attend conferences and working groups;
- Company outings, happy hours, hackathons, and tech talks;
- Receive a competitive compensation package with a strong benefits plan.
We are a Leader for Container Management in G2 (#2 after AWS)!
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search