Crystal Equation Corporation
Linkedin · Posted 3d ago
Senior System Engineer
Continue to application
Add your email once, then Caio opens the original posting.
Indexed description
What You'll Do
- Operate and scale Kubernetes platforms (EKS, CKS, GKE) across multiple cloud providers — managing cluster lifecycle, node pools, networking, and growth planning
- Provision HPC infrastructure through CI/CD systems across AWS, CoreWeave, GCP, and OCI
- Work with Slurm-based job scheduling to allocate GPU compute for AI training and inference workloads
- Build and maintain monitoring, alerting, and SLOs — contributing to operational excellence and incident response
- Collaborate daily with Networking, Storage, Security, and AI/ML platform teams
What We're Looking For
- 4+ years in infrastructure engineering, cloud platforms, or HPC
- Kubernetes experience (operating clusters, node pool management, upgrades)
- Terraform proficiency — writing and reviewing infrastructure-as-code daily
- Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre)
- Python for tooling and automation
- Slurm experience is a plus but not required
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search
Want help applying to roles like this?
Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search