Technology
Applications
About us
NewsContact us
← Back to open roles
Join the team

Foundation Models — World Model

Train and iterate the world models and World Action Models behind controllable prediction and embodied policy learning.

Responsibilities

  • Design, train, and continuously iterate world models for robotic manipulation, enabling controllable video and state prediction to support policy learning and planning.
  • Research and build World Action Models (WAMs) that understand and generate actions while modelling world dynamics, closing the loop between world prediction and action generation, and enabling an end-to-end transition from perception to decision-making.
  • Explore unified multimodal representations and pre-training strategies that integrate vision, language, action, audio, and other signals to build a transferable embodied foundation-model backbone.
  • Lead the design and optimisation of distributed training frameworks at the thousand-GPU scale, and systematically study the scaling laws of world models and WAMs in embodied intelligence.
  • Build evaluation benchmarks and systems for world models and action models, and establish the engineering pipeline from world model / WAM to policy / planning.

Qualifications

  • Master’s or Ph.D. degree in computer science, robotics, artificial intelligence, automation, or a related field. Publications in relevant areas at leading conferences such as CoRL, ICRA, NeurIPS, ICLR, CVPR, or RSS are preferred.
  • Familiarity with mainstream world-model and World Action Model methods, including DreamerV3, UniSim, Genie, JEPA, TD-MPC, and action-conditioned or action-predictive models, as well as their applications to control tasks.
  • In-depth research and engineering experience in at least one of the following areas. Experience across multiple areas is preferred, as these capabilities form the core foundation for world-model and WAM development:
    • Multimodal pre-training: familiarity with MLLMs, cross-modal alignment, and unified representations, including the CLIP and LLaVA families and multimodal diffusion models.
    • VLA model training: familiarity with mainstream paradigms and training workflows such as OpenVLA, the RT family, Octo, Diffusion Policy, and GR00T.
    • Video-generation or video-prediction pre-training: familiarity with diffusion-based or autoregressive video generation, Sora-, Genie-, or UniSim-style methods, and temporal-consistency modelling.
  • Experience in large-scale pre-training engineering, including participation in LLM, MLLM, or diffusion-model pre-training or training-framework development. Familiarity with Megatron, DeepSpeed, FSDP, or similar frameworks is required.
  • Strong engineering skills in Python and PyTorch. Experience with thousand-GPU-scale training and long-running training-stability optimisation is preferred.
  • Strong ability to reproduce research papers and translate algorithms into working systems. Representative open-source projects or competition awards are preferred.

Preferred Qualifications

  • First-author publications at leading conferences or representative open-source work in world models, video generation, or embodied intelligence.
  • Experience building a large-scale training cluster or training framework from the ground up.
  • Familiarity with the combination of reinforcement learning and world models, such as the Dreamer family or model-based reinforcement learning (MBRL), or experience with action generation or action-conditioned generation.

What We Offer

  • Substantial compute and data resources, together with research and deployment opportunities involving real robots.
  • The opportunity to explore the frontier of world models and WAMs with a top-tier team, with equal emphasis on research publications and engineering impact.
  • Compensation, equity, and benefits to be completed in accordance with company policy.
Apply by email

Location: Hangzhou, Beijing, Switzerland. Send your CV to info@awomo.ch. Suggested subject line: “Name + Position” — the apply button fills in the role for you, so just replace “Name”.