System Engineer
Indexed description
Headquartered in Silicon Valley, California, Penguin Solutions operates globally through a network of R&D, manufacturing, and sales locations. For nearly three decades, we have operated at the intersection of memory and AI/HPC infrastructure. That engineering expertise positions us to power the next generation of AI workloads, from training to inference and agentic AI at scale.
Penguin Solutions brings together differentiated infrastructure software, advanced memory, compute systems, end-to-end services, and industry-leading partner solutions in a full-stack AI factory platform designed to help customers deploy and scale AI workloads with speed and precision.
At Penguin Solutions, we value ideas over hierarchy and believe in servant leadership, where leaders enable teams to do their best work. We empower employees to take ownership, drive innovation, and grow through challenging work, continuous learning, and exposure to advanced AI tools and technologies. With flexibility where it matters and a strong focus on outcomes, Penguin Solutions is a place to do your best work, grow your career, and make a meaningful impact.
Job Overview
Penguin Solutions Managed Services provides dedicated, remote, Linux systems DevOps for complex, integrated environments involving high-performance computing, cloud, and enterprise systems. This position requires technical skills and the ability to understand, document, configure, administer, troubleshoot, and resolve issues in Linux environments. This is a customer-facing position.
Responsibilities
- Install, Deploy, Administer HPC Clusters
- Maintain, administer, and patch Linux Operating systems and associated software
- Work as part of a team to provide IT support and resolve errors.
- Analyze system log files and perform basic troubleshooting.
- Create Shell/Python/Ansible scripting
- Document processes through supporting System Engineers; follow and improve procedures to meet SLAs.
- Support users with Move/Add/Change requests.
- Troubleshoot errors and determine root cause.
- Respond to system alerts and monitoring, sometimes after hours
- Participate in weekly on-call rotation.
- Stay up-to-date on advancements Linux Operating Systems and associated software
- Bachelor's degree in Computer Science, Information Technology, or a related field; or equivalent experience.
- UNIX/Linux certification or equivalent experience.
- 7+ years of hands-on experience with UNIX/Linux server environments.
- HPC Systems Management knowledge.
- Linux systems administration skills and experience with open-source technologies.
- Understanding of Linux networking implementation and protocols.
- Ability to work in ITIL operating models.
- Able to install, configure, and tune software applications and provide overall support.
- Ability to speak English and communicate clearly and effectively with team members and clients.
- HPC: Application, Systems Management, OS, Optimization, Hardware and data center needs
- HPC/AI Performance Specialist and practical knowledge of the administration of High-Performance Computing (HPC) technologies, including cluster resource management, job scheduling, Ethernet networking, InfiniBand, etc.
- AI & Cloud: Virtualization, Applications, Container Orchestration, Systems Management, and Hardware design
- Data: High-Performance Storage and Parallel file systems used in HPC/AI and Cloud
- HPC cluster system admin experience.
- In-depth knowledge of Linux cluster technologies and optimization techniques.
- HPC Scheduler knowledge (SLURM, PBS, LSF).
- Will take initiative to refer to Application OEM/Vendor for Application operations, features, functions, and questions.
- Familiar with Supermicro hardware.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search