HPC Operations Engineer
Indexed description
We are looking for an adaptable hands-on individual, passionate about the details and nuances of managing Linux HPC environments at scale, and eager to tackle complex and unpredictable operational work as their primary job function.
What You'll Do:
- Provide front-line operational support for 24/7 Linux HPC compute, storage, and interconnects. Technologies involved include RDMA fabrics, parallel filesystems, HPC batch schedulers, FUSE filesystems, internal Jump software, multi-vendor hardware, cybersecurity requirements, a challenging and unpredictable client workload, and high user expectations
- Solve problem reports and questions posed by members of Jump's research community, escalating as needed and managing the entire problem lifecycle
- Respond to alerts in a timely fashion
- Participate in large, coordinated maintenance operations, including during evenings and weekends
- Work on global projects across a wide range of infrastructure
- Write code for diagnosing, resolving, and triaging difficult problems and automating frequently performed tasks
- Collaborate with team members and across teams to write code and testing infrastructures spanning both new and existing codebases in multiple programming languages
- Manage relationships with outside vendors, including traveling both domestically and internationally to meet with current and potential vendors
- Implement and support performance monitoring and fault monitoring systems
- Develop and improve systems and user documentation
- Develop and monitor the tools used to maintain a production computing environment
- Provide operational support as primary job function
- Adhere to all company cybersecurity and IT policies, including performing all work using only approved hardware and software
- Participate in an on-call rotation
- Other tasks as assigned or needed
- Work from company office an average of 5 days a week
- A desire for operational work as primary job function
- At least 2+ years of professional experience with Linux systems administration
- High performance computing (HPC), including parallel filesystems (e.g., Lustre, GPFS), batch systems (e.g., Slurm, Grid Engine), and high-performance network interconnects experience is a plus, but not required
- High proficiency with at least one programming/scripting language (e.g., Go, Python, C) and ability to learn additional languages quickly
- A compulsion to perform root cause analysis
- Strong verbal and written communication skills, including the ability to communicate effectively and efficiently with both coworkers and third-party vendors
- Strong collaboration skills with a willingness to undertake tasks of various technologies and complexities
- Ability to independently manage complex projects and multiple workstreams
- Strong sense of urgency
- Willingness to perform regular operational maintenance work during evenings and weekends and as needed
- Ability to work effectively in a busy, open floor plan office environment
- Reliable and predictable availability
- Private Medical, Vision and Dental Insurance
- Travel Medical Insurance
- Group Pension Scheme
- Group Life Assurance and Income Protection Schemes
- Paid Parental Leave
- Parking and Commuter Benefits
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search