Software Engineer, AI
Indexed description
About The Role
As a Reliability Engineer at Anthropic, you will play a crucial role in maintaining and enhancing the dependability of our AI systems, particularly focusing on Claude, our flagship large language model. The AI Reliability Engineering (AIRE) team collaborates across various departments to improve the robustness and resilience of our critical serving pathways—from SDKs and network layers to API infrastructure and hardware accelerators. Your work will involve designing and implementing systems that ensure high availability and low latency, managing incident responses, and supporting infrastructure that underpins our AI safety commitments. This position offers a unique opportunity to influence the reliability of cutting-edge AI systems at a company committed to safety and societal benefit, providing a dynamic and cross-disciplinary environment that values holistic system thinking.
Qualifications
- Bachelor’s degree or equivalent in a relevant field such as Computer Science, Engineering, or related disciplines
- Strong background in distributed systems, infrastructure, or reliability engineering
- Experience in operating large-scale model serving or training infrastructure, preferably with over 1000 GPUs
- Knowledge of ML hardware accelerators including GPUs, TPUs, or Trainium
- Understanding of ML-specific networking optimizations like RDMA and InfiniBand
- Familiarity with AI-specific observability tools and frameworks
- Experience with chaos engineering and resilience testing methodologies
- Contributions to open-source infrastructure or ML tooling are advantageous
- Excellent communication and collaboration skills with the ability to build strong cross-team relationships
- Demonstrated ownership and user-centric approach to system reliability
- Develop and define Service Level Objectives (SLOs) for large language model serving systems, balancing availability, latency, and development velocity
- Design, implement, and maintain monitoring and observability systems across the token processing pipeline
- Assist in designing and deploying high-availability serving infrastructure across multiple regions and cloud providers
- Lead incident response efforts for critical AI services, ensuring rapid resolution, conducting thorough incident reviews, and implementing systematic improvements
- Support the reliability and safety of safeguard model serving, ensuring alignment with safety commitments and operational excellence
- Collaborate with cross-functional teams to identify system vulnerabilities and implement resilience enhancements
- Contribute to the development of best practices for system reliability, scalability, and safety
- Competitive annual salary ranging from £325,000 to £390,000 GBP
- Comprehensive health and wellness benefits
- Opportunities for professional growth and development in a pioneering AI environment
- Flexible hybrid work policy with a minimum of 25% in-office presence
- Visa sponsorship available for eligible candidates
- Collaborative and innovative work culture focused on societal impact and safety
Anthropic is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate based on race, ethnicity, gender, sexual orientation, age, disability, or any other protected characteristic. We believe that diverse perspectives and backgrounds foster innovation and excellence, and we actively encourage candidates from all backgrounds to apply.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search