NVIDIA Storage Admin
Indexed description
Roles & Responsibilities
- Implementation
Establish namespace, OST/MDT layout, stripe/RAID policies; tiering for warm/cold datasets.
- Operations
Proactive detection of hot spots and metadata contention; schema for small?file handling.
- Performance & Optimization
- Reliability & Incident
Act as the storage SME during major incidents involving BeeGFS, Lustre, GPFS (IBM Spectrum Scale), NFS, NVMe-oF, Parallel File Systems, and Object Storage platforms.
Perform deep-dive troubleshooting and resolution of storage performance degradation, metadata bottlenecks, I/O latency spikes, filesystem corruption, capacity exhaustion, and hardware failures.
Conduct comprehensive Root Cause Analysis (RCA) for storage-related outages affecting GPU training, inferencing, AI pipelines, and high-performance computing (HPC) workloads.
Backup/DR for critical datasets; test restore time and RPO/RTO regularly; corruption and split?brain handling.
- Security & Compliance
Experience & Educational Requirement
BE/B-Tech or equivalent with Computer Science or Electronics & Communication
Certification- SNIA; vendor (NetApp/Dell/VAST) preferred.
Relevant Experience
- 7–12 years distributed storage; hands?on with Lustre/BeeGFS/Ceph, NVMe?oF, and S3 in GPU environments.
- Lustre/BeeGFS tooling, Ceph; GDS; Perf tools (fio, fs?digests); Prometheus/Grafana; Ansible.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search