May 2026
Kinematic Collapse and the Dyn-3D Benchmark. (a) Kinematic Collapse: Current MLLMs rely on superficial visual changes rather than genuine physical displacement. (b) Dyn-3D: Supervised by physical ground truth via Kinematic-GSPO, our model achieves explicit kinematic perception for accurate spatial and trajectory reasoning. (c) Performance: Our approach significantly outperforms standard baselines across diverse spatiotemporal tasks on the Dyn-3D benchmark.
Multimodal Large Language Models (MLLMs) are moving from static 2D perception toward dynamic 3D spatial understanding. In video, pixel changes mix scene geometry with camera motion. Without knowing the camera trajectory and motion magnitude, a model cannot tell whether parallax and scaling come from object depth or from fast ego-motion. Because monocular signals are scale-ambiguous, ego-motion perception is the metric bridge between the 2D visual stream and 3D space.
Current models often overfit to smooth trajectory priors rather than learning physical motion. Dense sampling and slow motion in mainstream video data create spurious correlations between visual change and displacement. When large displacements or non-smooth trajectories appear, spatial cognition fails. We call this failure Kinematic Collapse. In a pilot study of Qwen3.5-9B, increasing inter-frame displacement from 0.1 m to 5.0 m drops overall accuracy by 25.30 points (63.97% → 38.67%). Kinematic perception falls from 69.71% to 9.66%, and implicit trajectory reasoning stays below the 25% chance level.
We introduce the Dyn-3D Benchmark, the first evaluation of VLM ego-motion perception inside a 3D spatial framework. Using 3D Gaussian Splatting, we synthesize counterfactual video pairs with identical visual paths but different kinematics, totaling 16,063 test samples. We further release Dyn-3D-Instruct (33.6K samples) and propose TempoVista with Kinematic-GSPO, which injects metric physical ground truth into policy optimization. Experiments show that internalized kinematic perception improves motion estimation and robust spatial reasoning on Dyn-3D and VSI-Bench.
Two pilots test whether models understand physical motion or only superficial visual change.
As inter-frame displacement grows, local visual correspondences break. Spatial accuracy falls, kinematic perception collapses, and trajectory reasoning remains near chance. Models depend on the continuous visual cues of smooth training data.
On decoupled rotation and translation traps, Qwen3.5-9B scores 28.74% on pure rotation and 5.69% on pure translation (decoupled average 17.22%). Qwen3-VL-8B-Instruct and Qwen-VL-Max obtain 13.70% and 19.84%. Models routinely classify pure rotation as translation plus rotation, confusing viewpoint change with physical displacement.
Accuracy (%) on the 16,063-question Dyn-3D benchmark. Chance level is 25% for four-option questions.
| Subset | 0.1 m | 0.5 m | 1.0 m | 2.0 m | 5.0 m |
|---|---|---|---|---|---|
| All (#) | 2969 | 1463 | 3845 | 5037 | 2749 |
| All Acc. | 63.97 | 51.59 | 48.58 | 42.56 | 38.67 |
| Spatial Acc. | 67.53 | 62.90 | 57.08 | 51.84 | 52.76 |
| Kinematic Acc. | 69.71 | 30.71 | 31.29 | 23.04 | 9.66 |
| Trajectory Acc. | 0.00 | 3.33 | 17.24 | 13.47 | 4.91 |
Overview of the TempoVista framework. Left: kinematic-adaptive keyframe selection in a motion-aware SE(3) metric space, with denser sampling where kinematic variation is large. Right: Kinematic-GSPO generates candidate reasoning chains and scores them with a composite reward of spatial answer correctness, kinematic consistency, and format compliance.
Uniform sampling underrepresents sharp turns and fast translations. When camera extrinsics are available, TempoVista selects keyframes by greedy farthest-point sampling on SE(3), using a geodesic-inspired distance that balances translation and rotation. The selected frames form a geometric skeleton of the trajectory and preserve multi-object visibility.
After answer-supervised SFT, reinforcement learning asks the model for a structured motion trace—total path length, displacement magnitude, displacement vector, accumulated rotation, and average speed—plus the final answer. The reward is R = Rans + γ Rfmt + λ Rkin. Group-relative advantages and a clipped GSPO objective favor responses whose explicit motion interpretation matches scene-derived physical annotations.
| Benchmark | Modality | Videos | QAs | 3D Spatial | Ego-Motion | Counterfactual | Diagnostic Traps |
|---|---|---|---|---|---|---|---|
| ScanQA | Static 3D | — | 41K | — | — | — | — |
| SpatialVLM | Static 2D | — | 2B | ✓ | — | — | — |
| MVBench | Video | 4K | 4K | — | — | — | — |
| EgoSchema | Video | 5K | 5K | — | — | — | — |
| OpenEQA | Video | ~200 | 1.6K | ✓ | — | — | — |
| VSI-Bench | Video | 288 | 5K | ✓ | Partial | — | — |
| Dyn-3D (Ours) | Video | 835 | 16K | ✓ | ✓ | ✓ | ✓ |
Dyn-3D renders identical spatial paths with distinct dynamics (fast / smooth / slow) plus pure-rotation and pure-translation traps, so models cannot solve spatial questions from optical-flow shortcuts.
We reconstruct 447 indoor ScanNet++ scenes with 3D Gaussian Splatting (splatfacto in nerfstudio), then keep 443 eligible scenes: 263 for training, 167 held-out for evaluation, and 13 reserved. Each scene is rendered at 30 FPS under five motion types. All 16,063 evaluation questions are generated deterministically from restored metric 3D metadata and independently recomputed before scoring.
The five trajectory types are balanced (2,911–3,477 questions each). Diagnostic traps include 147 rotation and 21 translation items. Dyn-3D-Instruct provides 9,600 SFT samples and 24,000 RL samples, with no scene overlap against the benchmark.
Accuracy (%) with eight frames per video. TempoVista improves both InternVL-3.5-8B and Qwen3-VL-8B-Instruct; the latter reaches the best open-source overall score and surpasses Qwen-VL-Max by 11.4 points.
| Model | Kinematic | Spatial | Trajectory | Overall |
|---|---|---|---|---|
| Proprietary Models | ||||
| GPT-4o | 45.3 | 36.1 | 40.2 | 38.2 |
| Qwen-VL-Max | 53.2 | 46.2 | 71.9 | 49.8 |
| Open-Source Models | ||||
| LLaVA-Video | 40.7 | 28.4 | 61.8 | 33.7 |
| LLaVA-NeXT-Video | 19.1 | 30.5 | 34.3 | 28.7 |
| LLaVA-OneVision | 40.3 | 28.1 | 56.6 | 32.9 |
| Qwen2.5-VL-7B-Instruct | 30.7 | 37.4 | 44.3 | 36.7 |
| InternVL-3.5-8B | 52.0 | 47.8 | 70.4 | 50.6 |
| InternVL-3.5-8B + TempoVista | 64.5 (+12.5) | 52.7 (+4.9) | 77.7 (+7.3) | 57.1 (+6.5) |
| Qwen3-VL-8B-Instruct | 51.0 | 47.6 | 71.4 | 50.3 |
| Qwen3-VL-8B-Instruct + TempoVista | 62.6 (+11.6) | 58.3 (+10.7) | 81.8 (+10.4) | 61.2 (+10.9) |
Selected VSI-Bench multiple-choice spatial tasks with 32 frames per video. Accuracy (%).
| Model | Overall | Rel. Distance | Rel. Direction | App. Order | Route Plan |
|---|---|---|---|---|---|
| Proprietary Models | |||||
| GPT-4o | 36.1 | 37.0 | 41.3 | 28.5 | 31.5 |
| Gemini-1.5-Flash | 38.5 | 37.7 | 41.0 | 37.8 | 31.5 |
| Gemini-1.5-Pro | 44.0 | 51.3 | 46.3 | 34.6 | 36.0 |
| Open-Source Models | |||||
| Qwen2.5-VL-7B-Instruct | 34.5 | 38.0 | 37.4 | 28.0 | 28.4 |
| LLaVA-OneVision-7B | 34.1 | 42.5 | 35.2 | 24.4 | 29.4 |
| LLaVA-NeXT-Video-7B | 39.1 | 43.5 | 42.4 | 30.6 | 34.0 |
| InternVL-3.5-8B | 50.1 | 52.4 | 48.6 | 55.1 | 32.9 |
| InternVL-3.5-8B + TempoVista | 50.7 (+0.6) | 54.4 (+2.0) | 48.1 (-0.5) | 55.7 (+0.6) | 34.5 (+1.6) |
| Qwen3-VL-8B-Instruct | 51.3 | 53.5 | 47.9 | 60.2 | 32.0 |
| Qwen3-VL-8B-Instruct + TempoVista | 55.6 (+4.3) | 56.2 (+2.7) | 51.7 (+3.8) | 67.8 (+7.6) | 34.0 (+2.0) |
Overall Dyn-3D accuracy (%). Policy optimization provides most of the gain; the kinematic reward adds a further 1.1–1.6 points.
| Training Strategy | InternVL-3.5-8B | Qwen3-VL-8B-Instruct |
|---|---|---|
| Base model | 50.6 | 50.3 |
| SFT Only | 51.6 | 52.6 |
| Base GSPO | 55.5 | 60.1 |
| Kinematic-GSPO (Ours) | 57.1 | 61.2 |
Overall accuracy (%) under non-linear trajectories. Oracle SE(3) sampling outperforms uniform sampling by 1.5 points on both models; a visual flow proxy is worse than uniform sampling.
| Sampling Strategy | InternVL-3.5-8B + TempoVista | Qwen3-VL-8B-Instruct + TempoVista |
|---|---|---|
| Uniform | 55.6 | 59.7 |
| Flow Proxy | 55.1 | 59.0 |
| Oracle SE(3) FPS | 57.1 | 61.2 |
@misc{ding2026dyn3dunveilingresolvingegomotion,
title={Dyn-3D: Unveiling and Resolving Ego-Motion Ambiguity in Vision-Language Models},
author={Jiayu Ding and Zhuodong Liu and Lei Zhang and Manyu Xiong and Hongbo Jin and Haoran Tang and Hongbo Zhang and Changen Zhu and Wenbo Xing},
year={2026},
eprint={2609.01059},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.01059},
}