Back to search
Larsen & Toubro Linkedin · Posted 21d ago

NVIDIA Storage Admin

Mumbai

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Job Purpose

Provide high?throughput, consistent storage tiers (Scratch/HPS + Object) for large?scale training data ingest, checkpoints, and inference artifacts.

Roles & Responsibilities

  • Implementation

Design/expand Lustre/BeeGFS HPS; NVMe?oF and Object (S3) tiers; align with AI dataflow and GDS.

Establish namespace, OST/MDT layout, stripe/RAID policies; tiering for warm/cold datasets.

  • Operations

Capacity/performance planning; rebalance and failover testing; automate snapshots and checkpoint retention.

Proactive detection of hot spots and metadata contention; schema for small?file handling.

  • Performance & Optimization

Tune RDMA paths, page cache, IO schedulers; validate end?to?end I/O profiles for LLM training/inference.

  • Reliability & Incident

Lead P0/P1 critical incident response for enterprise-scale AI/ML storage infrastructure supporting NVIDIA GPU clusters, ensuring rapid service restoration and minimal impact to business-critical workloads.

Act as the storage SME during major incidents involving BeeGFS, Lustre, GPFS (IBM Spectrum Scale), NFS, NVMe-oF, Parallel File Systems, and Object Storage platforms.

Perform deep-dive troubleshooting and resolution of storage performance degradation, metadata bottlenecks, I/O latency spikes, filesystem corruption, capacity exhaustion, and hardware failures.

Conduct comprehensive Root Cause Analysis (RCA) for storage-related outages affecting GPU training, inferencing, AI pipelines, and high-performance computing (HPC) workloads.

Backup/DR for critical datasets; test restore time and RPO/RTO regularly; corruption and split?brain handling.

  • Security & Compliance

Multi?tenant encryption at rest and in transit, POSIX/ACLs, S3 policies, retention/legal holds.

Experience & Educational Requirement

BE/B-Tech or equivalent with Computer Science or Electronics & Communication

Certification- SNIA; vendor (NetApp/Dell/VAST) preferred.

Relevant Experience

  • 7–12 years distributed storage; hands?on with Lustre/BeeGFS/Ceph, NVMe?oF, and S3 in GPU environments.

Tools / Tech

  • Lustre/BeeGFS tooling, Ceph; GDS; Perf tools (fio, fs?digests); Prometheus/Grafana; Ansible.
Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search