Network Operations Center Technician II
Indexed description
ABOUT CIRRASCALE
Cirrascale Cloud Services provides high-performance cloud infrastructure purpose-built for deep learning, generative AI, and large-scale AI inference workloads. We specialize in dedicated GPU cloud solutions tailored to the unique needs of startups, research labs, and enterprise AI teams. Our mission is to accelerate AI innovation by combining powerful hardware with white-glove service and flexible, custom-built environments.
Key Responsibilities
- First-line whit glove response to alerts and incidents to systems and job failures
- Demonstrate understanding of GPU nodes and how they are deployed, networked, and clustered within a datacenter
- Assist customers with ticket triage and basic troubleshooting using the Jira (Atlassian) ticketing system
- Perform high-level monitoring and troubleshooting on all nodes and network equipment
- Remotely Troubleshoot Installed Servers & GPUs at various global datacenter locations
- Resolve complex and critical incidents within our datacenters
- Lead major incident response and coordination
- Review and optimize existing NOC procedures
- Work on capacity planning and performance monitoring
- Knowledge and experience working with Dell, SuperMicro & Lenovo type Servers is highly recommended
- Perform deep troubleshooting of GPU node failures, job preemption conflicts, and cluster imbalance
- Analyze alerts for GPU utilization inefficiencies, failed Machine Learning pipelines, or I/O bottlenecks
- Provide on-site and remote support to resolve urgent technical issues
- Document system configurations, updates, and inventory, maintaining accurate records of data center assets.
- Stay current with industry trends, emerging technologies, and best practices in HPC network operations center trends
Qualifications
- 2-4 years of experience in HPC, AI infrastructure, cloud systems, or related
- Understanding of scripting (Python, Bash, etc.), GPU resource monitoring preferred
- Solid understanding of HPC datacenter networking principles and experience with network troubleshooting
- Strong analytical and problem-solving skills, with the ability to work independently and manage multiple tasks
- Excellent communication skills and the ability to collaborate effectively with the customer and the team. Customer Service is a must
- Certifications: Advanced Linux or any other AI/ML certifications are a huge plus
- Experience in Datacenter Network Operations
- Experience with RMAs, logistics, shipping, and receiving a plus
- It is a plus with experience working in Jira (Atlassian) ticketing system.
- Proficient in Microsoft 365(Outlook, Word, Excel)
- Experience in Slack and Microsoft Teams
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search