Technology
Applications
About us
NewsContact us
← Back to the newsroom
Newsroom

StrucPhysVideo Achieves 45.5 on Physics-IQ with 30B-A3B Model

Research Release

Awomo introduces StrucPhysVideo, a family of video world models that combines physics-focused data curation with language- and action-conditioned prediction of how scenes evolve.

The approach starts with structured physical supervision. Its data pipeline combines motion-aware video segmentation, quality filtering and physical relevance checks. Structured captions distinguish camera motion from object behaviour and describe contact, deformation and state transitions over time, grounding training in observable physical events.

The text-image-to-video model uses a 30B-A3B (3B active parameters) sparse Mixture-of-Experts architecture. Training progressively emphasises physical dynamics while retaining general-domain video data. Experiments comparing captioning approaches across different model backbones further demonstrate the effectiveness of physics-focused supervision.

Figure 1. Physics-IQ Verified benchmark Results. StrucPhysVideo achieves the highest Physics-IQ Verified score among all compared models in the text-image-to-video setting. * denotes our reproduction, and results for all other models are taken from the benchmark snapshot dated 16 September 2026.
RankModelPhysics-IQ Verified
Input type: text-image-to-video
1StrucPhysVideo (Ours)45.5
2Cosmos3-Super-Image2Video42.7
3LingBot-Video*40.4
4MiniMax H3 (FL2VA)39.8
5Cosmos3-Nano37.3
6MiniMax H3 Max36.2
7Grok Imagine Video34.8
8Magi-1 24B + GeoPhys (BoN, OP)33.7
9Hunyuan Video 1.533.4
10Cosmos3-Edge32.7
11Wan 2.2 14B32.2
12CogVideoX-5B31.8
13Kandinsky-WM 1.030.8
14Magi-1 24B (OP)30.2
15Wan 2.2 5B27.7
16Sora 226.5
17P-Video25.3

On Physics-IQ, StrucPhysVideo achieves 45.5, the highest score among the models compared in the report’s text-image-to-video setting, surpassing Cosmos3-Super-Image2Video (64B) at 42.7. Results for all other models are from the benchmark as of 16 September 2026. The benchmark compares generated dynamics against real physical experiments, evaluating physical accuracy rather than visual realism alone.

Beyond text-image-to-video generation, the research extends to predicting visual outcomes from robot end-effector commands. Action conditioning, causal autoregressive generation and few-step distillation enable incremental robot video rollouts using four denoising steps. This shifts the focus from predicting what happens next to anticipating the consequences of a particular action.

Together, these contributions advance physical video prediction toward action-driven interaction—a foundation for Awomo’s longer-term work on planning and closed-loop robotic execution.

Further Information: