Data Center Operations System Engineer III (Chicago)
Indexed description
If you'd like to build the world's best AI cloud, join us.
- Note: This position requires presence in our Elk Grove, IL Data Center 5 days per week, 7/24 shift coverage
- Ensure new server, storage and network infrastructure is properly racked, labeled, cabled, and configured.
- Troubleshoot hardware and software issues in some of the world’s most advanced GPU and Networking systems.
- Document and update data center layout and network topology in DCIM software
- Work with supply chain & manufacturing teams to ensure timely deployment of systems and project plans for large-scale deployments
- Manage a parts depot inventory and track equipment through the delivery-store-stage-deploy-handoff process in each of our data centers
- Partner with HW Support teams to ensure data center hardware incidents with higher level troubleshooting challenges are resolved, reported on and solutions are disseminated to the large operations organization.
- Work with RMA team to ensure faulty parts are returned and replacements are ordered
- Follow installation standards and documentation for placement, labeling, and cabling to drive consistency and discoverability across all data centers
- Have strong past experiences with critical infrastructure systems supporting data centers, such as power distribution, air flow management, environmental monitoring, capacity planning, DCIM software, structured cabling, and cable management
- Be familiar with carrier DIA circuit test and turn ups, fiber testing and troubleshooting
- Basic knowledge of cable optics and the different types of use
- Solid understanding of single and three phase power theories
- PDU balancing and why it is important
- Familiar with multiple cable media types and their uses
- Knowledge of cold isle and hot isle containment
- Solid understanding of server hardware and boot process
- Ability to structure, collaborate and iteratively improve on complex maintenance MOPs.
- Working with product management, support, and other teams to align operational capabilities with company goals.
- Translating business priorities into technical and operational requirements.
- Supporting cross-functional projects where infrastructure plays a critical role.
- Founded in 2012, with 500+ employees, and growing fast
- Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
- We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
- Our values are publicly available: https://lambda.ai/careers
- We offer generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan that we all actually use
Compensation Range: $109K - $145K
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search