D2 Consulting
Linkedin · Posted 16d ago
HPC Infrastructure & Cluster Engineer
Continue to application
Add your email once, then Caio opens the original posting.
Indexed description
- ACTIVE TS/SCI SECURITY CLEARANCE REQUIRED**
Key Responsibilities:
- Cluster Administration: Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades
- Resource and Job Management: Configure, maintain, and optimize workload management and orchestration platforms, utilizing Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster
- Infrastructure Optimization: Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads
- Storage and Network Management: Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations
- Environment Configuration: Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment
- Security and Compliance: Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations
- 5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments
- Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand)
- Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g. Run:AI, SLURM)
- Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes
- Experience writing automation and configuration scripts (e.g. Bash, Python) to streamline cluster maintenance
- Proven ability to diagnose and resolve complex hardware, network, and OS-level issues
- Familiarity with parallel file systems and high-throughput storage architecture
- Prior experience engineering or managing high-speed GPU-to-GPU communication topologies
- All your information will be kept confidential according to EEO guidelines.
- Compensation is unique to each candidate and relative to the skills and experience they bring to the position. The salary range for this position is typically $170-$180k. This does not guarantee a specific salary as compensation is based upon multiple factors such as education, experience, certifications, and other requirements, and may fall outside of the above-stated range.
- Highlights of our benefits include Health/Dental/Vision, 401(k) match, Accrued PTO, STD/LTD/Life Insurance, Referral Bonuses, professional development reimbursement, and more!
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search
Want help applying to roles like this?
Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search