发表机构
Xiamen University; Shenzhen International Graduate School, Tsinghua University; China University of Mining and Technology; Shenzhen Technology University(厦门大学; 清华大学深圳国际研究生院; 中国矿业大学; 深圳技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对多模态大语言模型在连续动态场景感知中跟踪动态对象证据的局限,提出DynTrace框架,含DTV和DT-Token两个互补组件,能为模型提供基于几何视觉先验和结构化时空轨迹的动态对象证据,在多个基准测试中达先进水平。
AI 中文摘要
4D时空推理对于理解动态世界和实现具身交互至关重要。当前多模态大语言模型在静态场景理解和粗粒度4D任务中有强大能力,但在连续动态场景感知上有局限,尤其是跟踪动态对象证据进行连贯4D时空推理。受人类跟踪动态线索并补偿视角变化启发,提出DynTrace框架,含动态轨迹可视化(DTV)和动态跟踪令牌(DT-Token)。DTV将世界坐标轨迹投影到图像平面,DT-Token组织成动态跟踪图跟踪对象级动态线索等。该框架在多个基准测试中取得了领先成果,验证了跟踪动态对象证据对强大4D时空推理的重要性。
英文摘要
4D spatio-temporal reasoning, jointly modeling 3D spatial structure and temporal evolution, is essential for understanding dynamic worlds and enabling embodied interaction. While current Multimodal Large Language Models (MLLMs) show strong capabilities in static scene understanding and coarse-grained 4D tasks, they still have notable limitations in continuous dynamic scene perception, especially in tracking dynamic object evidence for coherent 4D spatio-temporal reasoning. This shortcoming stems mainly from relying on sparse frame-level observations, fragmenting continuous dynamic cues and leaving models unable to disentangle genuine object dynamics from camera-induced apparent motion. Inspired by humans tracking dynamic cues while compensating for viewpoint changes, we propose DynTrace, a training-free framework for 4D spatio-temporal reasoning with two complementary components. Dynamic Trajectory Visualization (DTV) reprojects world-coordinate trajectories onto the image plane, providing geometry-informed visual priors that disentangle genuine object dynamics from camera-induced apparent motion. Meanwhile, the Dynamic Trace Token (DT-Token), organized into a Dynamic Trace Graph (DTG), tracks object-level dynamic cues, trace evolution, and key moments, maintaining continuous dynamic object evidence for coherent 4D reasoning. Together, these two components equip MLLMs with continuously tracked dynamic object evidence, grounded in geometry-informed visual priors and structured spatio-temporal traces. DynTrace consistently improves open-source MLLMs, achieving state-of-the-art results on Dyn-Bench, VLM4D, and DSI-Bench, validating the importance of tracking dynamic object evidence for robust 4D spatio-temporal reasoning.
CommentsAccepted by ACM MM 2026