AI Infrastructure Engineer
Indexed description
Job Description
We are looking for an AI Infrastructure Engineer to join our global team responsible for managing and supporting large-scale AI infrastructure environments. You will help ensure the availability, performance, and operational stability of critical AI infrastructure platforms, working closely with a distributed team across regions to provide continuous coverage and support.
This is an opportunity to work hands-on with some of the most advanced Kubernetes-based AI infrastructure in production today, while contributing to the platforms and processes that keep it running reliably at scale.
Responsibilities:
- Manage and operate production AI infrastructure environments.
- Lead incident response and troubleshooting efforts and deliver timely service restoration during outages or performance degradations .
- Troubleshoot infrastructure and networking issues across bare-metal and/or cloud environments with multiple vendors.
- Conduct root cause analysis and drive product and operational improvements.
- Contribute and improve operational documentation and knowledge base.
- Collaborate with global team members across time zones to ensure continuous operational coverage, including occasional work during weekends and holidays.
- Proven experience managing and operating large-scale production systems (bare-metal and/or cloud).
- Solid working knowledge of Kubernetes with excellent, demonstrable troubleshooting skills.
- Experience configuring, customizing, and extending logging and monitoring tools (e.g., Prometheus, Grafana, ELK, or similar).
- Experience with infrastructure automation technologies and Infrastructure-as-Code practices (e.g., Ansible, Terraform, or similar).
- Effective verbal and written communication skills in English.
- Strong analytical and problem-solving skills, with the ability to work through complex, ambiguous technical issues.
- Willingness to occasionally work weekends and holidays.
- Previous experience building, scaling, and running High-Performance Computing (HPC) environments.
- Hands-on experiences with managing large scale Kubernetes platforms in production.
- A good understanding of NVIDIA GPU technologies and the associated software stack.
- Proficiency in scripting languages (e.g., Python, Bash, Go).
- Work with an established Silicon Valley leader in the cloud infrastructure industry;
- Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
- Be a part of cutting-edge, open-source innovation;
- Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
- Professional development and training;
- Attend conferences and working groups;
- Company outings, happy hours, hackathons, and tech talks;
- Receive a competitive compensation package with a strong benefits plan.
We are a Leader for Container Management in G2 (#2 after AWS)!
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search