Research Platform Engineer
Indexed description
We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.
Role Summary
We are seeking a talented and experienced platform engineer to join our Research Platform team. You'll work closely with our R&D team to build a cloud agnostic platform that improves the stability, scalability and velocity across the research department. You'll help create and optimize tools that empower researchers to make the most of high performance clusters.
Location: Paris / Warsaw / London (hybrid) or remote EU/UK with one hub visit per month.
What You Will Do
As a Platform Engineer, your responsibilities will include (but may not be limited to):
- Designing and implementing complex systems (e.g. scale our research CI with a strong focus toward reliability, reproducibility and speed)
- Building flexible yet solid and accessible development environment for researchers, so they can focus on core mission.
- Designing, implementing and advocating for solutions addressing large amounts of data and maintainable data pipelines.
- Optimizing a variety of builds: container images, large libraries compilation times, python environments...
- Building strong relationships with researchers, understanding their workflow and enabling them to achieve more by leveraging your expertise.
- Communicating and producing documentation or any content that will help them to make the most out of the tools and systems you'll build.
- Being part of the team that "platformizes" research and constantly improve the daily experience for researchers while avoiding future roadblocks.
- 5+ years of successful experience in a similar DX / DevOps / SRE role.
- Proficiency in software development (Python, Go...) and programming best practices.
- Exposure to site reliability engineering: root cause analysis, in-production troubleshooting, on-call rotations...)
- Exposure to infrastructure management: CI/CD, containerization, orchestration, infra-as-code, monitoring, logging, alerting, observability...).
- Technical product mindset (e.g. understanding how to debug poor adoption).
- Excellent problem-solving and communication skills (ability to contextualizing, gauging risks and getting buy-in for high stakes and impactful solutions).
- Ownership, high agency and constantly seeking to learn and improving things for others.
- Autonomous, self-driven and able to work well in a fast-paced startup environment.
- Low ego and team spirit mindset.
- First hand Bazel (or equivalent) experience.
- Strong knowledge of Python's ecosystem.
- Familiarity with GPU based workloads and ecosystems.
- Experience of full remote environments (you're comfortable with having some of your users on the other side of the globe).
- for the first week of their onboarding (accommodation and travelling covered)
- then at least 3 days per month
For the most up-to-date details on benefits available in your location, please refer to our Benefits page.
Privacy Policy
Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search