StrucPhysVideo Achieves 45.5 on Physics-IQ with 30B-A3B Model
Awomo introduces StrucPhysVideo, a family of video world models that combines physics-focused data curation with language- and action-conditioned prediction of how scenes evolve.
The approach starts with structured physical supervision. Its data pipeline combines motion-aware video segmentation, quality filtering and physical relevance checks. Structured captions distinguish camera motion from object behaviour and describe contact, deformation and state transitions over time, grounding training in observable physical events.
The text-image-to-video model uses a 30B-A3B (3B active parameters) sparse Mixture-of-Experts architecture. Training progressively emphasises physical dynamics while retaining general-domain video data. Experiments comparing captioning approaches across different model backbones further demonstrate the effectiveness of physics-focused supervision.
| Rank | Model | Physics-IQ Verified |
|---|---|---|
| Input type: text-image-to-video | ||
| 1 | StrucPhysVideo (Ours) | 45.5 |
| 2 | Cosmos3-Super-Image2Video | 42.7 |
| 3 | LingBot-Video* | 40.4 |
| 4 | MiniMax H3 (FL2VA) | 39.8 |
| 5 | Cosmos3-Nano | 37.3 |
| 6 | MiniMax H3 Max | 36.2 |
| 7 | Grok Imagine Video | 34.8 |
| 8 | Magi-1 24B + GeoPhys (BoN, OP) | 33.7 |
| 9 | Hunyuan Video 1.5 | 33.4 |
| 10 | Cosmos3-Edge | 32.7 |
| 11 | Wan 2.2 14B | 32.2 |
| 12 | CogVideoX-5B | 31.8 |
| 13 | Kandinsky-WM 1.0 | 30.8 |
| 14 | Magi-1 24B (OP) | 30.2 |
| 15 | Wan 2.2 5B | 27.7 |
| 16 | Sora 2 | 26.5 |
| 17 | P-Video | 25.3 |
On Physics-IQ, StrucPhysVideo achieves 45.5, the highest score among the models compared in the report’s text-image-to-video setting, surpassing Cosmos3-Super-Image2Video (64B) at 42.7. Results for all other models are from the benchmark as of 16 September 2026. The benchmark compares generated dynamics against real physical experiments, evaluating physical accuracy rather than visual realism alone.
Beyond text-image-to-video generation, the research extends to predicting visual outcomes from robot end-effector commands. Action conditioning, causal autoregressive generation and few-step distillation enable incremental robot video rollouts using four denoising steps. This shifts the focus from predicting what happens next to anticipating the consequences of a particular action.
Together, these contributions advance physical video prediction toward action-driven interaction—a foundation for Awomo’s longer-term work on planning and closed-loop robotic execution.