Join the team
Foundation Models — World Model
Train and iterate the world models and World Action Models behind controllable prediction and embodied policy learning.
Responsibilities
- Design, train, and continuously iterate world models for robotic manipulation, enabling controllable video and state prediction to support policy learning and planning.
- Research and build World Action Models (WAMs) that understand and generate actions while modelling world dynamics, closing the loop between world prediction and action generation, and enabling an end-to-end transition from perception to decision-making.
- Explore unified multimodal representations and pre-training strategies that integrate vision, language, action, audio, and other signals to build a transferable embodied foundation-model backbone.
- Lead the design and optimisation of distributed training frameworks at the thousand-GPU scale, and systematically study the scaling laws of world models and WAMs in embodied intelligence.
- Build evaluation benchmarks and systems for world models and action models, and establish the engineering pipeline from world model / WAM to policy / planning.
Qualifications
- Master’s or Ph.D. degree in computer science, robotics, artificial intelligence, automation, or a related field. Publications in relevant areas at leading conferences such as CoRL, ICRA, NeurIPS, ICLR, CVPR, or RSS are preferred.
- Familiarity with mainstream world-model and World Action Model methods, including DreamerV3, UniSim, Genie, JEPA, TD-MPC, and action-conditioned or action-predictive models, as well as their applications to control tasks.
- In-depth research and engineering experience in at least one of the following areas. Experience across multiple areas is preferred, as these capabilities form the core foundation for world-model and WAM development:
- Multimodal pre-training: familiarity with MLLMs, cross-modal alignment, and unified representations, including the CLIP and LLaVA families and multimodal diffusion models.
- VLA model training: familiarity with mainstream paradigms and training workflows such as OpenVLA, the RT family, Octo, Diffusion Policy, and GR00T.
- Video-generation or video-prediction pre-training: familiarity with diffusion-based or autoregressive video generation, Sora-, Genie-, or UniSim-style methods, and temporal-consistency modelling.
- Experience in large-scale pre-training engineering, including participation in LLM, MLLM, or diffusion-model pre-training or training-framework development. Familiarity with Megatron, DeepSpeed, FSDP, or similar frameworks is required.
- Strong engineering skills in Python and PyTorch. Experience with thousand-GPU-scale training and long-running training-stability optimisation is preferred.
- Strong ability to reproduce research papers and translate algorithms into working systems. Representative open-source projects or competition awards are preferred.
Preferred Qualifications
- First-author publications at leading conferences or representative open-source work in world models, video generation, or embodied intelligence.
- Experience building a large-scale training cluster or training framework from the ground up.
- Familiarity with the combination of reinforcement learning and world models, such as the Dreamer family or model-based reinforcement learning (MBRL), or experience with action generation or action-conditioned generation.
What We Offer
- Substantial compute and data resources, together with research and deployment opportunities involving real robots.
- The opportunity to explore the frontier of world models and WAMs with a top-tier team, with equal emphasis on research publications and engineering impact.
- Compensation, equity, and benefits to be completed in accordance with company policy.
Apply by email
Location: Hangzhou, Beijing, Switzerland. Send your CV to info@awomo.ch. Suggested subject line: “Name + Position” — the apply button fills in the role for you, so just replace “Name”.