HPC Data Storage Administrator (Research & HPC Data Platforms)
Indexed description
The Opportunity
Stanford’s Research Computing team is seeking a storage expert to join our Research & HPC Data Platforms group. This is a flexible-level posting; we are seeking an HPC Data Storage Administrator (Research & HPC Data Platforms) to maintain and expand our current world-class infrastructure.
You will work directly with our research and HPC data platforms team lead to manage a diverse environment of more than 100PB and 5 billion files, including high-speed Lustre, MinIO object storage, and Lustre HSM, among other platforms.
Why Stanford? You aren't just managing a storage cluster; you are a part of a data and storage ecosystem that supports Nobel-caliber research across all disciplines.
Storage Architect Key Responsibilities
- Architecture: Adapt and evolve the technology designs of existing systems to meet the needs of future computing platforms and research aims.
- Platform Management: Deliver on the scaling, reliability, security, compliance, operations, and lifecycle management of our primary research computing storage platforms, including for high-risk data.
- Tiered Storage Architecture: Oversee the integration of Lustre HSM on the Elm platform, managing data movement policies between parallel filesystems and MinIO object storage.
- Performance Engineering: Tune I/O for large-scale High Performance Computing and AI workloads.
- Community Stewardship: Represent Stanford within the Lustre community and other key community groups, contributing to the upstream roadmap and maintaining a vendor-neutral storage strategy.
- Education: Bachelor’s degree and eight years of relevant experience, or a combination of education and relevant experience.
- Expertise at Scale: 8+ years of hands-on experience architecting, building, and managing Lustre and ZFS or similar filesystems at the 20PB+ scale.
- Object Storage & HSM: Deep technical fluency in MinIO and Lustre HSM (copytools, policy engines like RobinHood) or similar tools.
- Kernel & Network Mastery: Expert-level knowledge of the Linux kernel and large-scale InfiniBand/Ethernet fabric tuning.
- In-depth Troubleshooting Experience: Must be capable of leading the debugging of issues such as kernel panics, LNet congestion, and metadata bottlenecks.
- Leadership: Proven experience mentoring junior admins and leading large-scale migration projects without data loss.
- Communication: Strong written and verbal communication skills.
- Platform Management: Contribute to the scaling, reliability, security, compliance, operations, and lifecycle management of our primary research computing storage platforms, including for high-risk data.
- Operational Excellence: Perform complex filesystem upgrades, kernel patches, and hardware refreshes with minimal downtime.
- Monitoring & Telemetry: In collaboration with others, build and maintain sophisticated observability stacks for real-time I/O tracking and trend analysis.
- User Support: Act as an escalation point for researchers struggling with complex I/O patterns, job failures, or data access issues.
- Maintenance: Manage the physical and logical health of the storage fleet, including RMA processes, firmware updates, and disk replacement cycles.
- Experience: 5+ years of Linux Systems Administration, with 3+ years specifically in an HPC or large-scale data environment.
- Technical Stack: Strong hands-on experience with Lustre, ZFS, MinIO, and/or similar technologies.
- Scripting: Advanced proficiency in scripting languages for automating routine storage tasks and parsing system logs.
- Hardware Mastery: Comfortable with the physical aspects of the role—diagnosing hardware failures and understanding power/cooling requirements for high-density storage.
- Communication: Strong written and verbal communication skills.
- Constantly perform desk-based computer tasks.
- Frequently sit, grasp lightly/fine manipulation.
- Occasionally stand/walk, writing by hand.
- Rarely use a telephone, lift/carry/push/pull objects that weigh up to 10 pounds.
- Consistent with its obligations under the law, the University will provide reasonable accommodations to applicants and employees with disabilities. Applicants requiring a reasonable accommodation for any part of the application or hiring process should contact Stanford University Human Resources by submitting a contact form .
- May work extended hours, evenings, and weekends.
- Interpersonal Skills: Demonstrates the ability to work well with Stanford colleagues and clients and with external organizations.
- Promote Culture of Safety: Demonstrates commitment to personal responsibility and value for safety; communicates safety concerns; uses and promotes safe behaviors based on training and lessons learned.
- Subject to and expected to stay in sync with all applicable University policies and procedures, including but not limited to the personnel policies and other policies found in Stanford's Administrative Guide, http://adminguide.stanford.edu.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search