Software Engineer (Product, Infrastructure and Platform Reliability)
Indexed description
Key Responsibilities
- Own reliability and SLA management across Sakana AI products, including Sakana Chat, Marlin, Fugu, and Namazu.
- Own production inference infrastructure for our LLM products, improving reliability, latency, throughput, GPU utilization, and cost efficiency.
- Operate LLM serving systems and support safe deployment and rollback workflows.
- Build monitoring, alerting, incident response, postmortems, and prevention practices, and participate in the on-call rotation.
- Translate enterprise security, availability, compliance, and SLA needs into infrastructure design.
- Support capacity planning and cost management, and provide technical requirements and capacity forecasts to the business team responsible for GPU and cloud procurement.
- Collaborate within a global team that works in both Japanese and English.
Required Qualifications
- Experience designing and operating production infrastructure in a cloud environment such as AWS or GCP.
- Experience operating containerized services, APIs, batch jobs, or model inference workloads in production.
- Experience managing and automating infrastructure with IaC or comparable tooling.
- Hands-on experience with monitoring, alerting, incident response, and postmortems.
- Ability to reason about GPU, cloud, and serving resource usage from cost, performance, and availability perspectives.
- Experience with LLM serving frameworks such as vLLM, TensorRT-LLM, or SGLang.
- English communication skills sufficient to drive technical projects with internal and external stakeholders, including discussions around requirements, cost, capacity, and priorities.
- Japanese language ability (if you are a native speaker or have passed JLPT N1/N2, please mention this in your application).
Preferred Qualifications
- Experience designing or operating production inference infrastructure for LLMs or machine learning models.
- Experience with Kubernetes, GKE, Vertex AI, Terraform, or comparable infrastructure tooling.
- Experience building or operating GPU clusters, especially NVIDIA H100/B200-class environments.
- Expertise in inference optimization techniques such as quantization, speculative decoding, and PD disaggregation.
- Knowledge of SRE practices, SLO design, capacity planning, or enterprise infrastructure.
Who We Are Looking For
- You can translate research and product needs into reliable production infrastructure.
- You optimize for latency, reliability, cost, operability, and product value.
- You are comfortable reading backend or inference pipeline code when needed.
- You can work effectively with business stakeholders on cost, capacity, and procurement decisions.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search