Join the team
Post-training
Own the pipeline from supervised fine-tuning and reward modelling to on-policy reinforcement learning on real robots.
Responsibilities
- Design supervised fine-tuning (SFT) data pipelines for robotic manipulation, and study instruction formats and multi-task data-mixture strategies.
- Build reward models that combine task-completion signals, physical plausibility, and alignment with language instructions.
- Implement on-policy reinforcement-learning fine-tuning using algorithms such as PPO, GRPO, and DPO, supporting both simulation and real-robot workflows.
- Research the application of reinforcement learning with verifiable rewards (RLVR) to embodied reasoning, improving the model’s planning capabilities for long-horizon tasks.
- Establish a standardised evaluation system covering manipulation accuracy, generalisation, safety, and related dimensions.
- Work closely with the pre-training, hardware, and data teams to build the complete pipeline from data to deployment.
Qualifications
- Master’s or Ph.D. degree in a relevant field. Publications at leading conferences such as CoRL, RSS, NeurIPS, or ICLR are preferred.
- Deep understanding of RLHF, PPO, DPO, and related algorithms, with hands-on experience building a complete post-training pipeline.
- Familiarity with VLA models such as OpenVLA, π0, and RoboFlamingo, or with LLM post-training.
- Experience training reinforcement-learning policies in simulation platforms such as Isaac or MuJoCo, with an understanding of sim-to-real challenges.
Apply by email
Location: Hangzhou, Beijing, Switzerland. Send your CV to info@awomo.ch. Suggested subject line: “Name + Position” — the apply button fills in the role for you, so just replace “Name”.