Storage Systems Engineer
Indexed description
- Responsibilities:
• Maintain the reliability, performance, and availability of enterprise storage and filesystem services.
• Operate storage platforms such as GPFS / IBM Storage Scale, NFS, Pure, VAST, NetApp, and NAS,
block, and object storage systems.
• Troubleshoot storage, filesystem, multipath, and storage client issues in production, including latency, contention, and capacity events.
• Monitor storage capacity, performance, and growth trends; respond to incidents and service degradations.
• Manage quotas, snapshots, replication, tiering, and data lifecycle policies.
• Perform production changes, firmware and software upgrades, data migrations, and remediation with strong operational discipline.
• Support backup, recovery, and disaster recovery processes and periodic restore testing.
• Develop scripts and automation tools using Python, Bash, or similar technologies.
• Work with hardware vendors, datacenter teams, and internal stakeholders to resolve issues quickly.
• Improve monitoring, alerting, runbooks, and operational processes.
• Participate in incident response and occasional on-call support.
- Mandatory Skills Description:
• 3+ years of experience in storage engineering, storage operations, or Linux infrastructure with significant storage responsibility.
• Good knowledge of storage technologies across file, block, and object storage.
• Solid Linux systems administration and production support skills.
• Understanding of storage client behavior, including NFS mounts, automount, multipath, and filesystem tuning.
• Understanding of networking fundamentals as they relate to storage, including DNS, throughput, and latency troubleshooting.
• Experience troubleshooting live issues using logs, metrics, and command-line tools.
• Scripting and automation experience with tools such as Python, Bash, Ansible, Salt, or Terraform.
• Strong communication, teamwork, and the ability to work effectively under pressure.
• Comfortable working close to production and solving real operational problems.
• Able to balance urgent incident response with long-term platform improvements.
• Organized, proactive, and careful with risk, change control, and communication.
- Nice-to-Have Skills Description:
• Deep experience with GPFS / IBM Storage Scale, ESS, VAST, Pure, NetApp, or similar platforms.
• Experience with SAN/NAS fabrics, RDMA or InfiniBand, and high-throughput data paths.
• Familiarity with backup, replication, snapshots, quotas, and lifecycle policies at scale.
• Exposure to cloud and hybrid storage, CI/CD, and DevOps workflows.
• Experience with Kubernetes storage integration such as CSI drivers and persistent volumes.
• Background supporting research, trading, HPC, or other high-availability environments.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search