Walden Robotics
Linkedin · Posted 12d ago
Senior Member of Technical Staff: Multimodal Pre-Training
Continue to application
Add your email once, then Caio opens the original posting.
Indexed description
Position SummaryYou'll be a senior technical leader for the pre-training of the omni-models at the core of our stack—single models trained across a wide range of objectives and modalities (LLM, VLM, video, and action) at massive scale. Working alongside our AI leadership and world-class team, you'll help set the research agenda, own high-impact architecture, objective, and data-mixture bets, and help push how far these base models scale.
Core Responsibilities
- Pre-Training Roadmap: Help drive the pre-training agenda for omni-models spanning language, vision-language, video, and action, in close partnership with the AI team.
- Omni-Model Objectives: Train a single model across a wide range of losses (LLM, VLM, video and image generation) at massive scale, and decide how to weight and balance them.
- Core Modeling: Lead modeling work across architecture, tokenization, and the scaling-law analysis behind compute allocation.
- Data & Infrastructure: Turn petabyte-scale text, image, and video data into training-ready mixtures with the data and infra teams.
- Technical Leadership: Help raise the bar on training reliability and evaluation, support and mentor teammates, and bring frontier ideas into our stack quickly.
- Large-Scale Pre-Training: Deep, hands-on experience pre-training large models—LLM, VLM, multimodal, or VLA—at significant scale, with concrete results.
- Multi-Objective Training: Experience training across multiple objectives or modalities—e.g. combining language, vision-language, and generative video/image losses—and balancing them in one model.
- Generative & World Models: Hands-on experience with generative modeling of images and/or video (diffusion, autoregressive, or masked objectives) and/or world-model approaches to physical dynamics.
- Scaling & Systems: Fluency with scaling laws, distributed training, and the practical failure modes of massive runs, and the ability to design around them.
- Technical Direction & Judgment: A track record of setting direction others build on, moving fluidly between research and production-grade code, and prioritizing ruthlessly under uncertainty.
- Experience pre-training LLMs or VLMs at scale.
- Experience training video-generation, image-generation, or world-model systems at scale.
- Experience with action/robotics modalities or VLA models for embodied agents.
- Contributions to widely used models, influential papers, or open-source training stacks.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search
Want help applying to roles like this?
Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search